LLMcompare

Search

Search for a command to run...

Meta·Llama / Muse·open·Llama 2

Code Llama 34B Instruct

Widely hosted Code Llama mid-size instruct — common Together / Fireworks coding baseline of 2023–2024.

  • Context 16K · #294/318
  • Blended $0.50 · #121/287
  • Speed 50 tok/s · #235/260
  • Benchmarks 2/49

Headline scores

Specs

Family
Llama / Muse
License
Llama 2
Open weights
Yes
Context
16K#294 of 318 in catalog
Max output
4K
Parameters
34B
Release
Aug 24, 2023
Cutoff
-
Modalities
text → text
Speed
50 tok/s · 0.45s TTFT
API
Together / Fireworks (ref.)

Benchmarks

Reasoning & knowledge

  • MMLU-Pro

    Harder multi-task language understanding across STEM and professional domains; reported results depend on prompt, shot count, and reasoning/tool settings. (source)

    —

    unreported

  • GPQA Diamond

    Graduate-level Google-proof Q&A in biology, physics, and chemistry (diamond subset); scores are only comparable when evaluation protocol and tool access match. (source)

    —

    unreported

  • Humanity's Last Exam

    Expert-level closed-ended questions across dozens of academic fields; tool access, modality, and evaluation protocol must match before scores are compared. (source)

    —

    unreported

  • AIME 2025

    American Invitational Mathematics Examination 2025 — contest math problems testing multi-step reasoning; model reports may differ by answer format and sampling setup. (source)

    —

    unreported

  • MATH-500

    500-problem subset of the MATH competition dataset covering algebra, geometry, and number theory; scores are protocol-sensitive and not interchangeable with AIME results. (source)

    —

    unreported

  • CritPt

    Complex Research using Integrated Thinking — Physics Test: 71 unpublished research-level physics challenges (decomposed into 190 checkpoint tasks) across 11 subfields, authored by 50+ working physicists. Reported here as Artificial Analysis's independently-run accuracy on the composite challenges; tool access materially changes results, so runs with and without code execution are not comparable. (source)

    —

    unreported

  • AA-LCR v1.1

    Artificial Analysis Long Context Reasoning: 100 open-answer questions requiring multi-step synthesis across documents of 10k-100k tokens (academic papers, financial reports, legal and government documents). Pass/fail graded by an LLM judge against official answers; scores are tied to the AA-LCR version and grader model and are not a needle-in-a-haystack retrieval measure. (source)

    —

    unreported

  • AA-Omniscience Accuracy

    Share of AA-Omniscience's 6,000 factual-recall questions answered correctly, across 42 topics in six domains (business, humanities and social sciences, STEM, health, law, and software engineering). This is the accuracy metric only — keep it distinct from the AA-Omniscience Index, a separate -100 to 100 metric that penalises hallucination and rewards abstention, and from the hallucination rate. (source)

    —

    unreported

  • MMMU-Pro

    Multimodal reasoning over 3,460 questions in six core disciplines. MMMU-Pro hardens the original MMMU in three ways: it filters out questions answerable by text-only models, expands the choice set from 4 to 10 options, and introduces a vision-only input format in which the question itself is embedded in a screenshot. Scores run well below MMMU and the two are not comparable. Artificial Analysis publishes one MMMU-Pro figure rather than separate standard and vision splits, so this value should not be read as either split on its own. (source)

    —

    unreported

  • IFBench

    Ai2's precise instruction-following benchmark: 58 verifiable out-of-domain output constraints held out from the training constraints, so it measures generalisation to unseen instructions rather than familiarity with common ones. Scored as constraint-satisfaction accuracy. (source)

    —

    unreported

  • Chartography

    Professional chart-understanding benchmark with 100 real-world chart questions across 12 domains; preserve tool access and scoring setup because tool-enabled and no-tool results are distinct. (source)

    —

    unreported

  • Chartography (With Tools)

    Chartography professional chart-understanding benchmark with tool access; keep separate from no-tool Chartography scores because tool use changes the evaluation setup. (source)

    —

    unreported

Math

  • MathArena

    MathArena contest-math leaderboard (AIME, AMC, olympiad-style); results depend on answer-format extraction and evaluation date. (source)

    —

    unreported

  • FrontierMath Tier 4 (v2)

    Epoch AI's FrontierMath benchmark, Tier 4 (v2 problem set): exceptionally difficult, original research-level mathematics problems vetted by expert mathematicians. Tier 4 is not comparable to FrontierMath Tiers 1-3, and v2 problem sets are not comparable to v1 results. (source)

    —

    unreported

Coding

  • SWE-bench Verified

    Resolves real GitHub issues in popular Python repositories (500-task human-verified subset); agent scaffold, patch budget, and test policy materially affect the result. (source)

    —

    unreported

  • SWE-bench Pro

    Long-horizon repository engineering across 1,865 tasks from 41 professional codebases; leaderboard results are agent-system scores, not model-only capability scores. (source)

    —

    unreported

  • SWE-bench Multilingual

    Real-world GitHub issue resolution across 300 tasks, 42 repositories, and 9 programming languages; agent scaffold and language mix matter. (source)

    —

    unreported

  • LiveCodeBench

    Contamination-resistant coding problems collected after model training cutoffs; score depends on the LiveCodeBench release and pass@k protocol. (source)

    —

    unreported

  • Terminal-Bench 2.1

    Terminal-Bench 2.1 contains 89 agentic terminal tasks. Scores are agent-system results and vary with harness, effort level, timeout, resource allocation, task revision, and evaluation date; they are not model-only measurements. (source)

    —

    unreported

  • Terminal-Bench 3

    Next-generation agentic terminal benchmark (Frontier-Bench); scores are agent-system results that vary with harness, effort, and evaluation date. (source)

    —

    unreported

  • Terminal-Bench 4.0

    Semantically-versioned continuation of Terminal-Bench with recalibrated task resources and a refreshed, less-saturated task set; scores are agent-system results that vary with harness, effort level, and evaluation date, and are not directly comparable to Terminal-Bench 2.x or 3 results. (source)

    —

    unreported

  • Aider Polyglot

    Multi-language coding agent benchmark: edit existing codebases to pass unit tests across languages; score depends on the Aider model prompt and edit strategy. (source)

    —

    unreported

  • SciCode

    Scientific research coding across physics, math, chemistry, biology, and materials — domain knowledge plus precise code synthesis; benchmark version and harness must match. (source)

    —

    unreported

  • CursorBench

    Cursor's IDE-native agent eval on ambiguous, multi-file tasks from real Cursor sessions — correctness under Cursor's agent scaffolding (v3.2), not a model-only test. (source)

    —

    unreported

  • SWE-Rebench

    Contamination-resistant software engineering: rolling window of fresh GitHub issues under a fixed ReAct agent scaffold; scores are tied to the published task window. (source)

    —

    unreported

  • NL2Repo-Bench

    Long-horizon repository generation: build a full installable Python library from a natural-language spec, graded by upstream pytest; harness and task release are part of the result. (source)

    —

    unreported

  • DeepSWE

    DataCurve's long-horizon software-engineering benchmark of 113 original tasks across TypeScript, Go, Python, JavaScript, and Rust with program-based verifiers; scores are agent-system results under the eval harness. (source)

    —

    unreported

  • WebDev Arena

    LMArena Code / WebDev human-preference Elo for front-end and agentic web development tasks; this is a time-varying leaderboard snapshot, not a fixed model property. (source)

    —

    unreported

  • BigCodeBench

    Realistic function-level coding benchmark with 1,140 tasks across 7 languages; pass@1 depends on the release and prompting. (source)

    29.0%

    #41 of 45

Tool use & function calling

  • Tau-bench

    Realistic tool-agent user interactions in airline, retail, and banking domains; scores are agent-system results, not model-only capability. (source)

    —

    unreported

  • Toolathlon-Verified

    Verified subset of Toolathlon, a tool-use benchmark measuring selecting, sequencing, and completing multi-step tasks with external tools; keep distinct from unverified Toolathlon runs. (source)

    —

    unreported

  • τ³-Banking

    97 fintech customer-support tasks (disputes, account freezes, provisional credits, product changes) where an agent must navigate roughly 700 interconnected policy documents (~195k tokens) and execute multi-step tool calls. Graded pass@1 against backend database state rather than conversational quality. This is an agent-system result and is distinct from τ²-bench Telecom and from earlier τ-bench releases. (source)

    —

    unreported

Agents & computer use

  • CyberGym

    Cybersecurity benchmark (UC Berkeley Sunblaze) evaluating AI agents on real-world vulnerability discovery and exploitation tasks; scores are agent-system results that vary with harness and environment. (source)

    —

    unreported

  • Agents' Last Exam

    Agent benchmark of long-horizon, real-world tasks requiring tool use and multi-step execution; results are agent-system measurements published by the evaluating provider. (source)

    —

    unreported

  • AutomationBench (Public)

    Zapier benchmark evaluating how reliably agents complete realistic business workflows across 47 simulated SaaS tools; scores are agent-system results and depend on the scaffold. (source)

    —

    unreported

  • AA-Briefcase v1.1

    Artificial Analysis's private agentic knowledge-work evaluation across four multi-week projects and 91 tasks. Its combined Elo metric includes rubric pass rate, analytical quality, and presentation quality; scores evaluate the agent system and its harness, not a model alone. (source)

    —

    unreported

  • Terminal-Bench-Science 0.1

    Agentic scientific-research workflows across the life, physical, Earth, mathematical, and engineering sciences, built on the Terminal-Bench harness with tasks authored by domain experts and verified programmatically. Scores are agent-system results with a wide per-model standard error (roughly ±3.5-4.5 pts) and are not comparable across benchmark releases. (source)

    —

    unreported

  • GDPval-AA v2

    Artificial Analysis's agentic evaluation of OpenAI's GDPval task set, covering real-world knowledge work across dozens of occupations and industries. Scored via blind pairwise comparison fit to a Bradley-Terry Elo model anchored at 1000 for human-expert deliverables; requires shell and browser access through an agent harness (Stirrup), so it is not a raw model-only capability score. (source)

    —

    unreported

  • OSWorld 2.0 (Partial Credit)

    OSWorld 2.0 long-horizon computer-use benchmark under partial-credit scoring, which awards credit for subtasks completed within a long real-world desktop workflow. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to pre-2.0 OSWorld results. (source)

    —

    unreported

  • OSWorld 2.0 (Strict)

    OSWorld 2.0 long-horizon computer-use benchmark under strict scoring, which requires the full task to complete for credit. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to partial-credit OSWorld 2.0 scores or to pre-2.0 OSWorld results. (source)

    —

    unreported

  • OSWorld 2.1 (Partial Credit)

    OSWorld-V2 2.1 task release with partial-credit scoring. Keep separate from OSWorld 2.0 and preserve the agent harness, task release, and tool setup in each score's provenance. (source)

    —

    unreported

Arena

  • LMArena Elo

    Blind pairwise human preference rating from the LMArena (Chatbot Arena) leaderboard; ratings drift with traffic, model aliases, sampling, and leaderboard methodology. (source)

    1,137

    #161 of 165

Composite indices

  • Artificial Analysis Intelligence Index v4.2

    Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.2 is a weighted average of ten independently-run evaluations across four categories: agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), general 30% (AA-Omniscience 15%, split 10% accuracy and 5% non-hallucination, GDP.pdf 10%, AA-LCR v1.1 5%), and scientific reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Index versions are not comparable to each other — v4.2 dropped GPQA Diamond and added AA-Briefcase and GDP.pdf relative to v4.1.1 — and several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted. (source)

    —

    unreported

  • Artificial Analysis Intelligence Index v4.3

    Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.3 is a weighted average of ten independently-run evaluations across four equally-weighted (25% each) categories: agents (AA-Briefcase, GDPval-AA v2, AutomationBench-AA), coding (Terminal-Bench 4.0, SciCode), general (AA-Omniscience, GDP.pdf, AA-LCR v1.1), and scientific reasoning (Humanity's Last Exam, CritPt). Relative to v4.2, v4.3 removes τ³-Banking, adds AutomationBench-AA, and swaps in Terminal-Bench 4.0 for Terminal-Bench 2.1. Index versions are not comparable to each other, so v4.3 scores must not be merged with v4.2 scores; several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted. (source)

    —

    unreported

  • Artificial Analysis Intelligence Index v4.3.2

    Artificial Analysis composite across AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Keep separate from prior index versions because its evaluation components changed. (source)

    —

    unreported

Open the compare picker