Helicone
Logs every model call by sitting in the request path, so nothing in your code has to be instrumented
What is Helicone?
Most tracing tools ask you to wrap your calls, which is fine until the code you need to see is inside a library, a background job or a service somebody else owns. Helicone takes the other route: point the base URL at it and every request is logged on the way through, with no SDK in the path. Because it is already in that position it does gateway work too — one key for many providers, automatic fallback when one is down, caching for repeated prompts. The trade is the position itself: a proxy is now between your application and the model, which is a component to run and a hop to account for. Agent runs group into sessions so a multi-step trace reads as one thing rather than thirty unrelated calls.
What can you do with Helicone?
- Log without changing the calling code — Changing the base URL is the whole integration, so calls made deep inside a library or another team's service are captured too.
- Read a multi-step run as one thing — Related calls group into a session, so an agent that made thirty requests to finish a task appears as one trace instead of thirty rows.
- Keep working when a provider does not — Because it is already in the request path, it can route across providers and fall back automatically when one starts failing.
- Stop paying twice for the same prompt — Repeated requests can be served from a cache, which matters in agent loops where the same context is resent on every turn.
- Turn logged traffic into a dataset — Recorded requests can be collected for evaluation or fine-tuning, so real usage becomes the test set rather than invented examples.
- Run it yourself if the data cannot leave — The platform is open source and self-hostable, which is the deciding factor when prompts contain material that must stay inside.
Before you choose Helicone
- A proxy in the request path is a component you now operate and a hop in every call — the reason to accept that is coverage you cannot get by instrumenting code you do not control.
- Self-hosting the full platform is a heavier proposition than adding a tracing library: it is a web application with its own datastores, not a package you import.
Frequently asked questions
Is Helicone free for commercial use?
Helicone is released under the Apache-2.0 licence — OSI-approved open source, which permits commercial use.
How can Helicone be deployed?
Helicone is available as Self-hosted / Managed cloud / Runs locally.
Documentation
Reproduced from the Helicone/helicone README, published under Apache-2.0. Read the original ↗
| 🔍 Observability | 🕸️ Agent Tracing | 🚂 LLM Routing |
|---|---|---|
| 💰 Cost & Latency Tracking | 📚 Datasets & Fine-tuning | 🎛️ Automatic Fallbacks |
Helicone is an AI Gateway & LLM Observability Platform for AI Engineers
- 🌐 AI Gateway: Access 100+ AI models with 1 API key through the OpenAI API with intelligent routing and automatic fallbacks. Get started in 2 minutes.
- 🔌 Quick integration: One-line of code to log all your requests from OpenAI, Anthropic, LangChain, Gemini, Vercel AI SDK, and more.
- 📊 Observe: Inspect and debug traces & sessions for agents, chatbots, document processing pipelines, and more
- 📈 Analyze: Track metrics like cost, latency, quality, and more. Export to PostHog in one-line for custom dashboards
- 🎮 Playground: Rapidly test and iterate on prompts, sessions and traces in our UI.
- 🧠 Prompt Management: Version prompts using production data. Deploy prompts through the AI Gateway without code changes. Your prompts remain under your control, always accessible.
- 🎛️ Fine-tune: Fine-tune with one of our fine-tuning partners: OpenPipe or Autonomi (more coming soon)
- 🛡️ Enterprise Ready: SOC 2 and GDPR compliant
🎁 Generous monthly free tier (10k requests/month) - No credit card required!
Quick Start ⚡️
-
Get your API key by signing up here and add credits at helicone.ai/credits
-
Update the
baseURLin your code and add your API key.import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://ai-gateway.helicone.ai", apiKey: process.env.HELICONE_API_KEY, }); const response = await client.chat.completions.create({ model: "gpt-4o-mini", // claude-sonnet-4, gemini-2.0-flash or any model from https://www.helicone.ai/models messages: [{ role: "user", content: "Hello!" }] }); -
🎉 You’re all set! View your logs at Helicone and access 100+ models through one API.
Self-Hosting Open Source LLM Observability
Docker
Helicone is simple to self-host and update. To get started locally, just use our docker-compose file.
# Clone the repository
git clone https://github.com/Helicone/helicone.git
cd docker
cp .env.example .env
# Start the services
./helicone-compose.sh helicone up
Helm
For Enterprise workloads, we also have a production-ready Helm chart available. To access, contact us at enterprise@helicone.ai.
Manual (Not Recommended)
Manual deployment is not recommended. Please use Docker or Helm. If you must, follow the instructions here.
Architecture
Helicone is comprised of five services:
- Web: Frontend Platform (NextJS)
- Worker: Proxy Logging (Cloudflare Workers)
- Jawn: Dedicated Server for serving collecting logs (Express + Tsoa)
- Supabase: Application Database and Auth
- ClickHouse: Analytics Database
- Minio: Object Storage for logs.
Integrations 🔌
Inference Providers
| Integration | Supports | Description |
|---|---|---|
| AI Gateway | JS/TS, Python, cURL | Unified API for 100+ providers with intelligent routing, automatic fallbacks, and unified observability |
| Async Logging (OpenLLMetry) | JS/TS, Python | Asynchronous logging for multiple LLM platforms |
| OpenAI | JS/TS, Python | Inference provider |
| Azure OpenAI | JS/TS, Python | Inference provider |
| Anthropic | JS/TS, Python | Inference provider |
| Ollama | JS/TS | Run and use large language models locally |
| AWS Bedrock | JS/TS | Inference provider |
| Gemini API | JS/TS | Inference provider |
| Gemini Vertex AI | JS/TS | Gemini models on Google Cloud’s Vertex AI |
| Vercel AI | JS/TS | AI SDK for building AI-powered applications |
| Anyscale | JS/TS, Python | Inference provider |
| TogetherAI | JS/TS, Python | Inference provider |
| Hyperbolic | JS/TS, Python | Inference provider |
| Groq | JS/TS, Python | High-performance models |
| DeepInfra | JS/TS, Python | Serverless AI inference for various models |
| Fireworks AI | JS/TS, Python | Fast inference API for open-source LLMs |
Frameworks
| Framework | Supports | Description |
|---|---|---|
| LangChain | JS/TS, Python | Use AI Gateway with LangChain for unified provider access |
| LlamaIndex | Python | Framework for building LLM-powered data applications |
| LangGraph | Python | Build stateful, multi-actor applications with LLMs |
| Vercel AI SDK | JS/TS | AI SDK for building AI-powered applications |
| Semantic Kernel | C#, Python | Microsoft’s AI orchestration framework |
| CrewAI | Python | Framework for orchestrating role-playing AI agents |
| ModelFusion | JS/TS | Abstraction layer for integrating AI models into JavaScript and TypeScript applications |
| PostHog | JS/TS, Python, cURL | Product analytics platform. Build custom dashboards. |
| RAGAS | Python | Evaluation framework for retrieval-augmented generation |
| Open WebUI | JS/TS | Web interface for interacting with local LLMs |
| MetaGPT | YAML | Multi-agent framework |
| Open Devin | Docker | AI software engineer |
| Mem0 EmbedChain | Python | Framework for building RAG applications |
| Dify | No code required | LLMOps platform for AI-native application development |
This list may be out of date. Don’t see your provider or framework? Check out the latest integrations in our docs. If not found there, request a new integration by contacting help@helicone.ai.
Additional Resources
-
LLM Cost API: We have the largest open-source API pricing database with 300+ models and providers such as OpenAI, Anthropic and more. Start querying here.
-
Data Management: Manage and export your Helicone data with our API or access it with our MCP server.
- Guides: ETL, Request Exporting
-
Data Ownership: Learn about Data Ownership and Autonomy
For more information, visit our documentation.