Source: https://www.tera.gw/inference
Retrieved: 2026-09-29T06:17:07.808Z (page retrieval, not inventory verification)
Content SHA-256: dfea47ceb3ca875bc7169ffb6ec736599a3e74fdda6c5f9e1fd0169a88ea3884

# Lowest-cost inference. Models you actually use.

Open-weight models: Qwen, Llama, DeepSeek, GLM, gpt-oss. OpenAI-compatible. US-processed, zero retention.

[Create your API key](<https://www.tera.gw/dashboard>) [Get free credits](<https://www.tera.gw/credits>)

[See what's inside ↓](<https://www.tera.gw/inference#why>)

Lowest cost

Per token. No minimums, no contracts.

Zero retention

Never trained on, never logged.

US-based

Team, GPUs, and data centers.

Drop-in API

Hermes Agent, OpenClaw, and any OpenAI-compatible tool.

Why Tera

01

### Lowest-cost open-source inference.

Per-token pricing on the models developers actually ship. No minimums, no contracts, no enterprise-tier negotiations. See the rate before you run the request.

02

### Your data stays yours.

Zero retention. Zero training. Zero human review. Prompts and completions are processed and discarded. Backed by a written privacy policy, not a press release.

03

### Processed entirely in the US.

US team. US-owned GPUs. US data centers. Your inference never crosses a border. Lower latency for US customers because the hardware is close to them.

04

### Open-weight only, by design.

We don't have a proprietary model to upsell you to. Tera is a pure execution layer for the open ecosystem. The day a better open model ships, it's on Tera.

Integrate

## Drop-in replacement for the OpenAI SDK.

Same API surface, lower-cost open-source models. Point your existing client at api.tera.gw and ship.

- 01

Get your API key

Sign up and grab a key from the dashboard.

- 02

Change one line

Set base\_url to api.tera.gw and keep the rest of your stack.

- 03

Pick a model

Qwen, Llama, DeepSeek, GLM, gpt-oss.

Python

TypeScript

Shell

api.tera.gw/v1

Copy

``````
from  openai  import  OpenAI

client  =  OpenAI(
    base_url = "https://api.tera.gw/v1" ,
    api_key = "tera_sk_..."
)

response  =  client.chat.completions.create(
    model = "openai/gpt-oss-120b" ,
    messages = [
        { "role" :  "user" ,  "content" :  "Hello, Tera." }
    ]
)

print (response.choices[ 0 ].message.content)
``````

Models

## Production-ready open models. Unified API.

Top-tier open-source models, served from owned U.S. infrastructure with zero data retention.

Anthropic

claude-fable-5

200K context

$10 in · $1 cache read · $50 out

Anthropic

claude-opus-4-8

200K context

$5 in · $0.5 cache read · $25 out

Anthropic

claude-sonnet-5

200K context

$2 in · $0.2 cache read · $10 out

DeepSeek

DeepSeek-R1-0528

164K context

$1.35 in · $5.4 out

DeepSeek

DeepSeek-V3.1

164K context

$0.6 in · $0.06 cache read · $1.7 out

DeepSeek

DeepSeek-V3.2

164K context

$0.56 in · $0.06 cache read · $1.68 out

DeepSeek

DeepSeek-V4-Flash

1M context

$0.19 in · $0.03 cache read · $0.51 out

DeepSeek

DeepSeek-V4-Pro

1M context

$1.74 in · $0.14 cache read · $3.48 out

Google

gemma-4-26B-A4B-it

262K context

$0.15 in · $0.01 cache read · $0.6 out

Meta

Llama-3.3-70B-Instruct

128K context

$0.72 in · $0.72 out

MiniMax

MiniMax-M2

197K context

$0.3 in · $0.03 cache read · $1.2 out

Moonshot

Kimi-K2.5

262K context

$0.6 in · $0.1 cache read · $3 out

Moonshot

Kimi-K2.6

262K context

$0.95 in · $0.16 cache read · $4 out

Moonshot

kimi-k2-thinking

262K context

$0.6 in · $0.06 cache read · $2.5 out

Moonshot AI

Kimi-K3

1M context

$3 in · $0.3 cache read · $15 out

OpenAI

gpt-5.5

1M context

$5 in · $0.5 cache read · $30 out

OpenAI

gpt-5.6-luna

1M context

$0.2 in · $0.02 cache read · $1.2 out

OpenAI

gpt-5.6-sol

1M context

$5 in · $0.5 cache read · $30 out

OpenAI

gpt-5.6-terra

1M context

$2 in · $0.2 cache read · $12 out

OpenAI

gpt-oss-120b

131K context

$0.09 in · $0.36 out

OpenAI

gpt-oss-20b

131K context

$0.07 in · $0.25 out

Qwen / Alibaba

Qwen3-235B-A22B-Instruct-2507

262K context

$0.22 in · $0.88 out

Qwen / Alibaba

Qwen3-Coder-480B-A35B-Instruct

262K context

$0.22 in · $0.02 cache read · $1.8 out

Qwen / Alibaba

Qwen3-Next-80B-A3B-Instruct

262K context

$0.15 in · $1.2 out

Qwen / Alibaba

Qwen3-Next-80B-A3B-Thinking

262K context

$0.15 in · $1.2 out

Z.ai

GLM-4.7

200K context

$0.6 in · $2.2 out

Z.ai

GLM-5

200K context

$1 in · $0.1 cache read · $3.2 out

Z.ai

GLM-5.2

262K context

$1.49 in · $0.27 cache read · $4.62 out

Full model list and pricing at [/pricing](<https://www.tera.gw/pricing>).

Built for agent workloads

## Everything agent builders actually need.

Designed for the inference patterns that agents generate: streaming, tool calls, reasoning traces, and high call volume.

OpenAI-compatible tool calling

The tools API works without modification. Any framework that uses OpenAI function calling routes through Tera unchanged.

Streaming responses

Server-sent events on every model. Low time-to-first-token so your agents stay responsive under load.

Reasoning content

Chain-of-thought traces surfaced as a dedicated field on reasoning models. Same API shape, no custom parsing.

Not billed for failed requests

5xx errors and rate-limit responses (429) are not charged. You pay only for tokens that are successfully processed.

No platform fee

Per-token pricing with no minimums, no contracts, and no monthly base charge. See the rate before you run the request.

Works with Hermes Agent, OpenClaw, Cline, and any OpenAI-compatible framework.

FAQ

## Common questions.

### Is Tera's API OpenAI-compatible?

Yes. Change one line: set base\_url (or baseURL) to https://api.tera.gw/v1. The rest of your OpenAI SDK code works unchanged. Chat completions, streaming, and tool calling are all supported.

### What models are available?

Qwen, Llama, DeepSeek, GLM, gpt-oss, MiniMax, and Gemma, with more added as production-quality open models ship. The full list with per-token pricing is at tera.gw/pricing.

### How do I get started?

Sign up at tera.gw/dashboard, create an API key, and you can make your first call within minutes. No approval process, no contracts, no minimums.

### What does Tera cost?

Per-token pricing, no platform fee, no minimums. Prices vary by model. See the full rate card at tera.gw/pricing. Failed requests (5xx errors, 429 rate limits) are not billed.

### Where is my data processed?

All inference runs on US-owned infrastructure in US data centers. Prompts and completions are processed and discarded. Zero logging, zero retention, zero training on your data.

### Does Tera support streaming and tool calling?

Yes. Set stream: true for server-sent events on any model. The OpenAI tools API (function calling) is supported. Reasoning models also surface chain-of-thought traces in a dedicated reasoning field.

### What agent frameworks work with Tera?

Any framework that speaks the OpenAI API works without modification: LangChain, LlamaIndex, AutoGen, CrewAI, and tools like Hermes Agent, OpenClaw, and Cline.

Self-serve

## Direct API access. Live now.

Sign up, get a key, and start running open-weight models today. No waitlist, no approval process. Pay per token, no minimums.

[Create your API key](<https://www.tera.gw/dashboard>) [Already have an account? Sign in →](<https://www.tera.gw/dashboard>)

---
Interactive controls above are described, not submitted. Reading does not reserve capacity or authorize payment.
For GPU requirements: https://www.tera.gw/gpu-rentals#capacity-brief
