The non-thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It gives faster direct answers across text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, and long-context work.
Added Jul 15, 2026
Model weightsContext Window
1.0M
Max Output
32.8K
Input Price (Auto)
$1.00/1M
Output Price (Auto)
$4.05/1M
Cache Read (Auto)
$0.17/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
42.3
Coding Index
52.1
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
87.2%
Better than 83% of models compared
HLE
Humanity's Last Exam
31.9%
Better than 83% of models compared
AA-LCR
Long context reasoning evaluation
73.3%
Better than 74% of models compared
Coding
SciCode
Python programming for scientific computing
46.1%
Better than 31% of models compared
Last updated Aug 17, 2026
Artificial AnalysisProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare Inkling with similar models from the same provider or model family.
Inkling Thinking
thinkingmachines/inkling:thinkingThe thinking version of Thinking Machines' 975B-parameter open-weights Mixture-of-Experts generalist with 41B active parameters. It reasons natively over text, images, and audio, and is built for agentic coding, tool use, detailed instruction following, long-context work, and controllable thinking effort.
Inkling Small
thinkingmachines/Inkling-SmallThe direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.
Inkling Small Thinking
thinkingmachines/Inkling-Small:thinkingThe reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.
Qwen3.8 Max 0902
alibaba/qwen3.8-max-0902Qwen3.8 Max 0902 is Alibaba's September 2 checkpoint of its flagship Qwen3.8 Max model for coding, knowledge work, data analysis, and long-running agent workflows. It supports text, image, video, PDF input, selectable thinking, tool calling, structured output, and a near-million-token context window.
Gemini 3.8 Flash
google/gemini-3.8-flashGoogle's fast multimodal model for agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Its capabilities, limits, reasoning behavior, and pricing currently mirror Gemini 3.7 Flash.
Muse Spark 1.3
meta/muse-spark-1.3Meta's Muse Spark 1.3 is a frontier multimodal reasoning model for long-horizon coding and agentic workflows, with strong gains in computer use, browsing, professional tool use, codebase understanding, instruction following, and million-token retrieval. It accepts text, images, audio, video, and files, supports tool calling and structured output, and always reasons before answering.