GitHub hosts HunyuanVideo, Tencent's open-source framework for large-scale video generation models, enabling AI-driven video creation.
Inference
About Inference
Inference provides a single endpoint to route requests across multiple AI model providers including OpenAI, Anthropic, Google, Mistral, and others. The platform allows explicit model selection while maintaining scoped API keys, prepaid credits, and live cost tracking in one dashboard. Users can manage rate limits, view request timelines, and monitor spend across chat, responses, image generation, and embeddings endpoints. The service supports OpenAI-compatible SDKs, enabling developers to switch models without rewriting integration code. A public catalog lists available models with pricing, endpoint types, and code samples for evaluation before account creation. The dashboard tracks credits, rate limits, and model-level costs, with plans that scale from free trials to production tiers supporting higher request volumes and multiple modalities.
Key features
- Unified API for chat, responses, images, and embeddings
- Explicit model selection with fallback routing
- Scoped API keys with per-key limits
- Live cost and latency tracking
- OpenAI-compatible SDK integration
- Request timeline and usage history
- Prepaid credits with monthly caps
- Public model catalog with pricing
Use cases
- Routing AI model requests across multiple providers
- Controlling spend and usage for AI-powered applications
- Testing and comparing different model providers
Pros
- Single endpoint for multiple model providers
- Explicit model selection with fallback routing
- Scoped API keys with per-key limits and usage tracking
- OpenAI-compatible SDK integration
- Live cost and latency visibility
Cons
- Requires account creation for full functionality
- Limited to supported model providers
- No free tier beyond basic trial credits
Frequently asked questions about Inference
Can I view models before creating an account?
Yes. The public catalog shows available models, their pricing, endpoint types, and code samples for evaluation without requiring an account.
Will my existing SDK code work with Inference?
Yes. Inference supports OpenAI-compatible SDKs, so you can switch to its base URL and pass the model ID in your existing code without rewriting integrations.
How do API keys work in Inference?
You can create scoped keys for specific apps, teams, or environments, with usage limits and tracking applied per key while keeping credentials contained.
Where can I track spend and limits?
The dashboard provides real-time visibility into credits, rate limits, request history, and model-level costs after account creation.
Does Inference support multiple AI model providers?
Yes. It routes requests across providers like OpenAI, Anthropic, Google, Mistral, Meta, and others through a single endpoint.
Can I use Inference for image generation and embeddings?
Yes. The platform supports chat, responses, image generation, and embeddings endpoints under a unified dashboard and API structure.