Zhipu AI
FP8
Function Calling
CN/EN

GLM-5.2-FP8

Enterprise-grade foundation model served at FP8 — optimized for high-throughput structured JSON extraction and function calling.

Context window
1M
Modalities
Text
Reasoning
Yes
Vision
No
Pooled throughput
5M TPM
Price
$0.25 in · $0.80 out
# request this model on one TATC key model="glm-5-2-fp8"

Connect GLM-5.2-FP8 in minutes

One TATC key, one Base URL — swap glm-5-2-fp8 in the model field of any OpenAI-compatible client.

  • Seamless OpenAI/LangChain integration
  • Single account, consolidated billing
  • Instant key rotation & spend limits
from openai import OpenAI

client = OpenAI(
    base_url="https://api.tatc.cloud/v1",  # Unified API endpoint
    api_key="sk-xxxxxxxx",                 # One key for all supported models
)

# GLM-5.2-FP8
response = client.chat.completions.create(
    model="glm-5-2-fp8",                   # Easily switch to any model ID
    messages=[{"role": "user", "content": "Hello TATC"}],
)
Platform capabilities

Every model ships with the full platform

Six pillars that keep multi-model traffic fast, secure and auditable — identical for every model in the catalogue.

Multi-Model Gateway

Unified access to leading multi-modal foundation models behind a single OpenAI-compatible endpoint. Switch models per request with zero client re-configuration.

Ultra-Low Latency Routing

Edge-optimized traffic management with per-model pool tuning and 24/7 P95 latency monitoring, delivering optimized Time-To-First-Token (TTFT) for production workloads.

Enterprise Security & Compliance

Per-tenant key isolation, VPC peering, request-level audit logging, and zero-data-retention (ZDR) architecture aligned with international data protection and regional privacy frameworks.

Usage Analytics & FinOps

Granular per-model and per-key usage dashboards tracking tokens, cost, latency, and success rates — fully exportable for chargeback and enterprise FinOps.

Flexible Key Management

Issue scoped API keys per team, project, or environment with hard spend caps, TPM limits, and instant rotation. Full data export with zero platform lock-in.

Universal SDK Compatibility

Drop-in compatibility with standard OpenAI, LangChain, and LiteLLM SDKs. Point your base URL to our gateway and keep your existing client codebase intact.

Start with GLM-5.2-FP8 today

Talk to our team for a Token API key — or request a dedicated capacity quote for production workloads.