Gemini 1.5 Pro

Name: Gemini 1.5 Pro
Brand: Google
Price: 1.25 USD
Rating: 44.7 (12 reviews)

Closed

Google

Proprietary

text

vision

Gemini 1.5Released 2y ago

Avg score

44.7

/ 100

Context

2.0M

Output limit

Input price

$1.25 /M

Output price

$5.00 /M

Pricing verified 1y ago · Doubles to $2.50/$10.00 above 128k tokens

Benchmarks

preference

Chatbot Arena EloFresh

Elo

Crowdsourced pairwise human preference rankings of LLM responses. Higher Elo means more frequently preferred by users.

math

AIME 2024High risk

American Invitational Mathematics Examination 2024 problems. Three-digit integer answers; very hard for non-reasoning models.

OTIS Mock AIME 2024-2025Fresh

AIME-style competition problems written specifically for the OTIS mock contest, then run as an evaluation by Epoch AI. Closer in spirit to the public AIME but with novel problems unlikely to appear in training data.

coding

HumanEvalSaturated

% pass@1

164 hand-written Python programming problems scored by passing unit tests. Saturated for frontier models.

vision

MMMUSome risk

Massive Multi-discipline Multimodal Understanding; college-exam level questions with images across 30+ subjects.

MathVistaSome risk

Math reasoning over visual contexts (charts, figures, geometry).

long context

RULER 128kFresh

Long-context retrieval and reasoning suite. We report the 128k token effective-context score.

performance

Output SpeedN/A

tok/s

Median sustained output speed in tokens per second on the model's first-party API for medium-length prompts. Higher is faster.

Time to First TokenN/A

Median time from request to first output chunk in milliseconds on the model's first-party API for medium-length prompts. Lower is snappier; reasoning models are penalised here because they think before talking.

reasoning

Humanity's Last ExamFresh

A challenging multi-disciplinary exam aggregating expert-written questions from across academic fields. Designed to discriminate at the very top of the capability range when MMLU-style tests saturate.

ARC-AGI 2Fresh

Second-generation ARC challenge testing fluid reasoning over abstract visual puzzles. Resists training-data memorisation by construction: each puzzle is novel and solutions require multi-step pattern induction. Frontier models are only just starting to score above chance on the harder tier.

composite

Frontier CompositeFresh

ECI

Saturation-resistant composite capability score stitched together from ~40 underlying benchmarks using Item Response Theory. Each benchmark is weighted by its fitted difficulty and discriminative slope, so doing well on hard, contamination-resistant evals (FrontierMath, ARC-AGI 2, Humanity's Last Exam) moves the score and saturated benchmarks contribute almost nothing. Imported per-model from Epoch AI's published index; we anchor it to the same min-max scale we use for every other benchmark so it's directly weightable in scenarios.

Reliability monitor

Loading drift signal…

Hosted endpoints

No third-party hosts tracked for this model — available only from its primary provider.

Compare with...

vs GPT-4o vs GPT-4o mini vs o1 vs o1-mini vs o3 vs o4-mini vs o3-mini vs GPT-4 Turbo vs GPT-4.1 vs GPT-5