Private AI
Browse and discover the best AI language models for conversations, coding, and creative writing.
Google's fast multimodal model for agentic workloads, including coding, tool use, image understanding, PDF and document extraction, audio, and video. Its capabilities, limits, reasoning behavior, and pricing currently mirror Gemini 3.7 Flash.
Features
Context
1.0M
Max Output
65.5K
Date Added
Sep 2, 2026
Input:
$0.75/1M
Output:
$3.75/1M
Cache:
Read $0.08/1M · Write $0.08/1M
Est./msg:
$0.0026
Not included in subscription
Meta's Muse Spark 1.3 is a frontier multimodal reasoning model for long-horizon coding and agentic workflows, with strong gains in computer use, browsing, professional tool use, codebase understanding, instruction following, and million-token retrieval. It accepts text, images, audio, video, and files, supports tool calling and structured output, and always reasons before answering.
Features
Context
1.0M
Max Output
943.7K
Date Added
Sep 2, 2026
Input:
$1.25/1M
Output:
$4.25/1M
Cache:
Read $0.15/1M
Est./msg:
$0.0034
Not included in subscription
Meta's Muse Spark 1.3 Contributor is a frontier multimodal reasoning model for long-horizon coding and agentic workflows, with strong gains in computer use, browsing, professional tool use, codebase understanding, and million-token retrieval. It accepts text, images, audio, video, and files, supports tool calling and structured output, and always reasons before answering. Prompts and outputs may be used by Meta for training and to improve its products.
Qwen3.8 Max 0902 is Alibaba's September 2 checkpoint of its flagship Qwen3.8 Max model for coding, knowledge work, data analysis, and long-running agent workflows. It supports text, image, video, PDF input, selectable thinking, tool calling, structured output, and a near-million-token context window.
Claude Fable 5.1 improves on Fable 5 across agentic coding, long-running workflows, front-end and visual code generation, finance, analysis, and knowledge work, with more concise plans and summaries. Anthropic retains prompts and outputs for 30 days; Zero Data Retention is not available.
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Abliteration.ai's default unrestricted large text reasoning model is derived from GLM-5.3 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
GLM-5.3 is Z.AI's open-weight reasoning model for complex software engineering, autonomous agents, vulnerability research, and long-horizon tasks. This text-only TEE deployment is verified through the selected provider: Redpill attestation with signed completion receipts or the official Tinfoil SDK's ATC/EHBP verification.
IBM Granite 4.2 8B is an Apache 2.0-licensed dense model with native step-by-step reasoning and specialized training for agentic work. It can plan before acting, sequence tools, navigate codebases, work in terminals, and verify results across coding, search, mathematics, science, and complex instruction-following tasks.
GLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters. This TEE deployment is verified through the selected provider: Redpill attestation with signed completion receipts or the official Tinfoil SDK's ATC/EHBP verification.
Hy4 Preview is Tencent's 770B-parameter mixture-of-experts model with 49B active parameters. It is designed for coding agents, complex tool-use workflows, and productivity tasks, with a 1M-token context window and configurable reasoning effort.
Abliteration.ai's unrestricted multimodal reasoning model is weight-modified to reduce refusals and answer prompts the original model might reject. It supports text and image input, structured output, automatic prompt caching, and a 262K-token context window.
Abliteration.ai's unrestricted large text reasoning model is derived from GLM-5.2 and weight-modified to reduce refusals compared with the base model. It supports native tool calling, structured output, automatic prompt caching, and a one-million-token context window.
ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
This is the same underlying model as Qwen3.8 Max, exposed under its architecture-based 2.4T A95B name for easier discovery. It uses the identical routing, pricing, capabilities, and non-thinking mode.
Qwen3.8 Flash is Alibaba's latest fast multimodal model, with a million-token context window for coding, agentic workflows, visual understanding, long documents, codebases, and videos.
Qwen3.8 27B is an open-weight dense vision-language model from Alibaba for reasoning, coding, professional workflows, multimodal interaction, tool use, and structured output. Running inside a TEE (Trusted Execution Environment), with provider attestation support.
An experimental vision-enabled DeepSeek V4 Flash model that adds image understanding while retaining the text, reasoning, coding, tool-calling, and agent capabilities of the base model.