Tencent·Other·open·Apache 2.0
Hy4 Preview
Tencent Hy Team's Aug 28, 2026 open-weight flagship (770B MoE / 49B active, Apache 2.0) with a 1M context window and default high reasoning effort. Text-only instruct weights on Hugging Face; hosted API at $0.834/$2.501 per 1M tokens via TokenHub and OpenRouter.
- Context 1M · #25/318
- Blended $1.25 · #176/287
- Benchmarks 7/49
- Thinking high · no_think / high
Headline scores
Specs
- Family
- Other
- License
- Apache 2.0
- Open weights
- Yes
- Context
- 1M#25 of 318 in catalog
- Max output
- -
- Parameters
- 770B (49B act.)
- Release
- Aug 28, 2026
- Cutoff
- -
- Modalities
- text → text
- Speed
- —
- API
- Tencent Cloud TokenHub
- Thinking
- no_think / highreasoning_effort; catalog uses high
Benchmarks
reasoning_effort · highest high · source
List price does not change by thinking level. Thinking tokens bill as output. Levels without a recorded score stay on the measured level.
Reasoning & knowledge
MMLU-Pro
Harder multi-task language understanding across STEM and professional domains; reported results depend on prompt, shot count, and reasoning/tool settings. (source)
—
unreported
GPQA Diamond
Graduate-level Google-proof Q&A in biology, physics, and chemistry (diamond subset); scores are only comparable when evaluation protocol and tool access match. (source)
92.3%
#20 of 215
Humanity's Last Exam
Expert-level closed-ended questions across dozens of academic fields; tool access, modality, and evaluation protocol must match before scores are compared. (source)
—
unreported
AIME 2025
American Invitational Mathematics Examination 2025 — contest math problems testing multi-step reasoning; model reports may differ by answer format and sampling setup. (source)
—
unreported
MATH-500
500-problem subset of the MATH competition dataset covering algebra, geometry, and number theory; scores are protocol-sensitive and not interchangeable with AIME results. (source)
—
unreported
CritPt
Complex Research using Integrated Thinking — Physics Test: 71 unpublished research-level physics challenges (decomposed into 190 checkpoint tasks) across 11 subfields, authored by 50+ working physicists. Reported here as Artificial Analysis's independently-run accuracy on the composite challenges; tool access materially changes results, so runs with and without code execution are not comparable. (source)
—
unreported
AA-LCR v1.1
Artificial Analysis Long Context Reasoning: 100 open-answer questions requiring multi-step synthesis across documents of 10k-100k tokens (academic papers, financial reports, legal and government documents). Pass/fail graded by an LLM judge against official answers; scores are tied to the AA-LCR version and grader model and are not a needle-in-a-haystack retrieval measure. (source)
—
unreported
AA-Omniscience Accuracy
Share of AA-Omniscience's 6,000 factual-recall questions answered correctly, across 42 topics in six domains (business, humanities and social sciences, STEM, health, law, and software engineering). This is the accuracy metric only — keep it distinct from the AA-Omniscience Index, a separate -100 to 100 metric that penalises hallucination and rewards abstention, and from the hallucination rate. (source)
—
unreported
MMMU-Pro
Multimodal reasoning over 3,460 questions in six core disciplines. MMMU-Pro hardens the original MMMU in three ways: it filters out questions answerable by text-only models, expands the choice set from 4 to 10 options, and introduces a vision-only input format in which the question itself is embedded in a screenshot. Scores run well below MMMU and the two are not comparable. Artificial Analysis publishes one MMMU-Pro figure rather than separate standard and vision splits, so this value should not be read as either split on its own. (source)
—
unreported
IFBench
Ai2's precise instruction-following benchmark: 58 verifiable out-of-domain output constraints held out from the training constraints, so it measures generalisation to unseen instructions rather than familiarity with common ones. Scored as constraint-satisfaction accuracy. (source)
—
unreported
Chartography
Professional chart-understanding benchmark with 100 real-world chart questions across 12 domains; preserve tool access and scoring setup because tool-enabled and no-tool results are distinct. (source)
—
unreported
Chartography (With Tools)
Chartography professional chart-understanding benchmark with tool access; keep separate from no-tool Chartography scores because tool use changes the evaluation setup. (source)
—
unreported
Math
MathArena
MathArena contest-math leaderboard (AIME, AMC, olympiad-style); results depend on answer-format extraction and evaluation date. (source)
—
unreported
FrontierMath Tier 4 (v2)
Epoch AI's FrontierMath benchmark, Tier 4 (v2 problem set): exceptionally difficult, original research-level mathematics problems vetted by expert mathematicians. Tier 4 is not comparable to FrontierMath Tiers 1-3, and v2 problem sets are not comparable to v1 results. (source)
—
unreported
Coding
SWE-bench Verified
Resolves real GitHub issues in popular Python repositories (500-task human-verified subset); agent scaffold, patch budget, and test policy materially affect the result. (source)
—
unreported
SWE-bench Pro
Long-horizon repository engineering across 1,865 tasks from 41 professional codebases; leaderboard results are agent-system scores, not model-only capability scores. (source)
65.7%
#5 of 40
SWE-bench Multilingual
Real-world GitHub issue resolution across 300 tasks, 42 repositories, and 9 programming languages; agent scaffold and language mix matter. (source)
82.9%
#1 of 21
LiveCodeBench
Contamination-resistant coding problems collected after model training cutoffs; score depends on the LiveCodeBench release and pass@k protocol. (source)
—
unreported
Terminal-Bench 2.1
Terminal-Bench 2.1 contains 89 agentic terminal tasks. Scores are agent-system results and vary with harness, effort level, timeout, resource allocation, task revision, and evaluation date; they are not model-only measurements. (source)
85.4%
#11 of 110
Terminal-Bench 3
Next-generation agentic terminal benchmark (Frontier-Bench); scores are agent-system results that vary with harness, effort, and evaluation date. (source)
—
unreported
Terminal-Bench 4.0
Semantically-versioned continuation of Terminal-Bench with recalibrated task resources and a refreshed, less-saturated task set; scores are agent-system results that vary with harness, effort level, and evaluation date, and are not directly comparable to Terminal-Bench 2.x or 3 results. (source)
—
unreported
Aider Polyglot
Multi-language coding agent benchmark: edit existing codebases to pass unit tests across languages; score depends on the Aider model prompt and edit strategy. (source)
—
unreported
SciCode
Scientific research coding across physics, math, chemistry, biology, and materials — domain knowledge plus precise code synthesis; benchmark version and harness must match. (source)
—
unreported
CursorBench
Cursor's IDE-native agent eval on ambiguous, multi-file tasks from real Cursor sessions — correctness under Cursor's agent scaffolding (v3.2), not a model-only test. (source)
—
unreported
SWE-Rebench
Contamination-resistant software engineering: rolling window of fresh GitHub issues under a fixed ReAct agent scaffold; scores are tied to the published task window. (source)
—
unreported
NL2Repo-Bench
Long-horizon repository generation: build a full installable Python library from a natural-language spec, graded by upstream pytest; harness and task release are part of the result. (source)
—
unreported
DeepSWE
DataCurve's long-horizon software-engineering benchmark of 113 original tasks across TypeScript, Go, Python, JavaScript, and Rust with program-based verifiers; scores are agent-system results under the eval harness. (source)
64.3%
#16 of 29
WebDev Arena
LMArena Code / WebDev human-preference Elo for front-end and agentic web development tasks; this is a time-varying leaderboard snapshot, not a fixed model property. (source)
1,624
#8 of 74
BigCodeBench
Realistic function-level coding benchmark with 1,140 tasks across 7 languages; pass@1 depends on the release and prompting. (source)
—
unreported
Tool use & function calling
Tau-bench
Realistic tool-agent user interactions in airline, retail, and banking domains; scores are agent-system results, not model-only capability. (source)
—
unreported
Toolathlon-Verified
Verified subset of Toolathlon, a tool-use benchmark measuring selecting, sequencing, and completing multi-step tasks with external tools; keep distinct from unverified Toolathlon runs. (source)
74.1%
#6 of 27
τ³-Banking
97 fintech customer-support tasks (disputes, account freezes, provisional credits, product changes) where an agent must navigate roughly 700 interconnected policy documents (~195k tokens) and execute multi-step tool calls. Graded pass@1 against backend database state rather than conversational quality. This is an agent-system result and is distinct from τ²-bench Telecom and from earlier τ-bench releases. (source)
—
unreported
Agents & computer use
CyberGym
Cybersecurity benchmark (UC Berkeley Sunblaze) evaluating AI agents on real-world vulnerability discovery and exploitation tasks; scores are agent-system results that vary with harness and environment. (source)
—
unreported
Agents' Last Exam
Agent benchmark of long-horizon, real-world tasks requiring tool use and multi-step execution; results are agent-system measurements published by the evaluating provider. (source)
—
unreported
AutomationBench (Public)
Zapier benchmark evaluating how reliably agents complete realistic business workflows across 47 simulated SaaS tools; scores are agent-system results and depend on the scaffold. (source)
—
unreported
AA-Briefcase v1.1
Artificial Analysis's private agentic knowledge-work evaluation across four multi-week projects and 91 tasks. Its combined Elo metric includes rubric pass rate, analytical quality, and presentation quality; scores evaluate the agent system and its harness, not a model alone. (source)
—
unreported
Terminal-Bench-Science 0.1
Agentic scientific-research workflows across the life, physical, Earth, mathematical, and engineering sciences, built on the Terminal-Bench harness with tasks authored by domain experts and verified programmatically. Scores are agent-system results with a wide per-model standard error (roughly ±3.5-4.5 pts) and are not comparable across benchmark releases. (source)
—
unreported
GDPval-AA v2
Artificial Analysis's agentic evaluation of OpenAI's GDPval task set, covering real-world knowledge work across dozens of occupations and industries. Scored via blind pairwise comparison fit to a Bradley-Terry Elo model anchored at 1000 for human-expert deliverables; requires shell and browser access through an agent harness (Stirrup), so it is not a raw model-only capability score. (source)
—
unreported
OSWorld 2.0 (Partial Credit)
OSWorld 2.0 long-horizon computer-use benchmark under partial-credit scoring, which awards credit for subtasks completed within a long real-world desktop workflow. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to pre-2.0 OSWorld results. (source)
—
unreported
OSWorld 2.0 (Strict)
OSWorld 2.0 long-horizon computer-use benchmark under strict scoring, which requires the full task to complete for credit. Agent-system result tied to harness, tool access, and the specific task release used; not comparable to partial-credit OSWorld 2.0 scores or to pre-2.0 OSWorld results. (source)
—
unreported
OSWorld 2.1 (Partial Credit)
OSWorld-V2 2.1 task release with partial-credit scoring. Keep separate from OSWorld 2.0 and preserve the agent harness, task release, and tool setup in each score's provenance. (source)
—
unreported
Arena
LMArena Elo
Blind pairwise human preference rating from the LMArena (Chatbot Arena) leaderboard; ratings drift with traffic, model aliases, sampling, and leaderboard methodology. (source)
—
unreported
Composite indices
Artificial Analysis Intelligence Index v4.2
Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.2 is a weighted average of ten independently-run evaluations across four categories: agents 30% (AA-Briefcase 15%, GDPval-AA v2 10%, τ³-Banking 5%), coding 20% (Terminal-Bench 2.1 10%, SciCode 10%), general 30% (AA-Omniscience 15%, split 10% accuracy and 5% non-hallucination, GDP.pdf 10%, AA-LCR v1.1 5%), and scientific reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Index versions are not comparable to each other — v4.2 dropped GPQA Diamond and added AA-Briefcase and GDP.pdf relative to v4.1.1 — and several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted. (source)
—
unreported
Artificial Analysis Intelligence Index v4.3
Artificial Analysis's own composite metric on a 0-100 scale, not a percentage. v4.3 is a weighted average of ten independently-run evaluations across four equally-weighted (25% each) categories: agents (AA-Briefcase, GDPval-AA v2, AutomationBench-AA), coding (Terminal-Bench 4.0, SciCode), general (AA-Omniscience, GDP.pdf, AA-LCR v1.1), and scientific reasoning (Humanity's Last Exam, CritPt). Relative to v4.2, v4.3 removes τ³-Banking, adds AutomationBench-AA, and swaps in Terminal-Bench 4.0 for Terminal-Bench 2.1. Index versions are not comparable to each other, so v4.3 scores must not be merged with v4.2 scores; several components are agent-harness evaluations, so the index is not a model-only capability measurement. Only listings Artificial Analysis has fully measured are recorded here; its estimated index values are omitted. (source)
—
unreported
Artificial Analysis Intelligence Index v4.3.2
Artificial Analysis composite across AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. Keep separate from prior index versions because its evaluation components changed. (source)
—
unreported