Chat, embedding, reasoning, and other model APIs — not only LLMs
Claude models with tool use, the MCP connector, Agent Skills, and the Claude Agent SDK
Run models at the edge from Workers, with an OpenAI-compatible endpoint and AI Gateway
Gemini multimodal models via Google AI Studio and Vertex AI, with function calling and MCP support
Serverless and dedicated inference for open models, plus a hosted MCP endpoint over Hub tools
Open-source proxy and SDK exposing 100+ model providers behind one OpenAI-compatible API
GPT-5 series models, embeddings, and the Responses API with native MCP and tool use
One OpenAI-compatible endpoint in front of hundreds of hosted models, with automatic fallback and per-key spend limits
Run thousands of open-source models — LLM, image, video, audio — behind one API
Unified model endpoint with failover, spend limits, and observability, wired into the AI SDK
Dedicated deployments and autoscaling inference for open and custom models
Wafer-scale inference API serving open models at very high tokens per second
Models tuned for enterprise RAG, reranking, search, and classification
High-quality machine translation and document translation over REST
Low-cost reasoning and chat models with an OpenAI-compatible interface
Fast, cheap serving for open-source LLMs, vision, and audio models
Ultra-fast LLM inference on custom LPU hardware, tuned for low time-to-first-token
Open and commercial models with fast inference, function calling, and OCR endpoints
Search-grounded LLM API returning cited answers over live web content
AI gateway with conditional routing, guardrails, caching, and governance across model providers
Run, fine-tune, and serve open-source models via API
Same query as JSON: /api/v1/services?category=thinkAdd to Claude Code:claude mcp add botfriendly …