Chat, embedding, reasoning, and other model APIs — not only LLMs
Run models at the edge from Workers, with an OpenAI-compatible endpoint and AI Gateway
One OpenAI-compatible endpoint in front of hundreds of hosted models, with automatic fallback and per-key spend limits
Run thousands of open-source models — LLM, image, video, audio — behind one API
Unified model endpoint with failover, spend limits, and observability, wired into the AI SDK
Dedicated deployments and autoscaling inference for open and custom models
Models tuned for enterprise RAG, reranking, search, and classification
Open and commercial models with fast inference, function calling, and OCR endpoints
Search-grounded LLM API returning cited answers over live web content
AI gateway with conditional routing, guardrails, caching, and governance across model providers
Run, fine-tune, and serve open-source models via API
Hosted agentic AI assistant: OpenAI-compatible API, BYOK 50+ providers, free ache/* + paid catalog.
Resold OpenAI-compatible model API. Prepaid tokens; whole-cent rounding up per successful request, US$0.01 minimum.
Same query as JSON: /api/v1/services?category=think&has_llms_txt=trueAdd to Claude Code:claude mcp add botfriendly …