Models

Every model. One runtime.

XR is model-agnostic. Use frontier models, open-weight, or your own fine-tunes. Route by cost, latency, or capability.

Recommended
XR

XR Core 1

Our in-house agentic model, tuned for tool use and long-horizon tasks.

Context
1M tokens
Speed
Fastest
Tag
Flagship
Available
All plans
Use this model
Recommended
Anthropic

Claude Opus 4.5

Frontier reasoning for complex planning, code generation, and deep analysis.

Context
2M tokens
Speed
Balanced
Tag
Reasoning
Available
All plans
Use this model
OpenAI

GPT-5

Versatile frontier model with strong generalist performance.

Context
1M tokens
Speed
Fast
Tag
General
Available
All plans
Use this model
Google

Gemini 2.5 Pro

Strong multimodal reasoning across text, code, images, and video.

Context
1M tokens
Speed
Fast
Tag
Multimodal
Available
All plans
Use this model
DeepSeek

DeepSeek V3

High-performance open-weight coding model.

Context
128K tokens
Speed
Very Fast
Tag
Open-weight
Available
All plans
Use this model
Meta

Llama 4 Maverick

Open-weight model with native multimodal capabilities.

Context
1M tokens
Speed
Fast
Tag
Open-weight
Available
All plans
Use this model
Alibaba

Qwen 3 Max

Top-tier multilingual open-weight model.

Context
1M tokens
Speed
Fast
Tag
Open-weight
Available
All plans
Use this model
Groq

Groq Llama 405B

405B Llama served on Groq's LPU for sub-300ms latency.

Context
128K tokens
Speed
Realtime
Tag
Ultra-low latency
Available
All plans
Use this model

Bring your own model.

Self-host Ollama, vLLM, Llama.cpp, or any OpenAI-compatible endpoint. XR’s model gateway handles routing, caching, and cost controls out of the box.