Bitcoin.com AI logo
Provider logo

Mercury Coder Small

mercury-coder-small
Provider logo

Mercury Coder Small

mercury-coder-small

Model by Inception AI. A diffusion large language model that runs incredibly quickly (500+ tokens/second) while matching Claude 3.5 Haiku and GPT-4o-mini. 1st in speed on Copilot arena, and matching 2nd in quality.

Context Window

32.8K

Max Output

16.4K

Input Price (Auto)

$0.25/1M

Output Price (Auto)

$1.00/1M

Cache Read (Auto)

$0.13/1M

Benchmarks

Performance metrics and benchmarks

No benchmark data is available yet for this model.

Providers

Auto routing is available for this model. Explicit provider selection is not available.

Loading provider options…

Compare Mercury Coder Small with similar models from the same provider or model family.

Mercury 2.5 Preview

inception/mercury-2.5-preview

Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.

Mercury 2

mercury-2

Inception Labs' fastest reasoning model with tool calling and structured outputs support.

Inkling Small

thinkingmachines/Inkling-Small

The direct-answer version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It accepts text and images, and is designed for coding, tool use, instruction following, and general conversational work.

Inkling Small Thinking

thinkingmachines/Inkling-Small:thinking

The reasoning version of Thinking Machines' 276B-parameter open-weights multimodal Mixture-of-Experts model with 12B active parameters. It reasons over text and images, and is designed for agentic coding, tool use, instruction following, and long workflows with controllable thinking effort.

Mistral Small 4 119B Thinking

mistralai/mistral-small-4-119b-2603:thinking

Mistral Small 4 with reasoning enabled (reasoning_effort=high). A hybrid MoE model with deep step-by-step reasoning for complex prompts, coding, and multi-step problem solving.

Qwen 3 Coder 480B

qwen/qwen3-coder

Qwen 3 Coder 480B, a 480 billion total parameter model with 35B active, and 160 total experts with 8 active. Performs similar to Claude 4 Sonnet in coding benchmarks, but does so at a much lower price.