Helicone AI Gateway: A Complete Guide with Practical Examples

Complete Guide to the Helicone AI Gateway (with Practical Examples)

Table of Contents

What Is Helicone AI Gateway?

Helicone AI Gateway is an open-source AI gateway, giving you access to 100+ AI models from most LLM providers. Instead of managing separate integrations for OpenAI, Anthropic, Google, and others, you use one consistent interface and API key to reach all of them.

Unlike traditional API gateways, Helicone AI Gateway includes zero markup pricing and has built-in observability by default, so every request is automatically logged, tracked, and analyzed without additional configuration or pricing.

Problems Helicone AI Gateway Solves

Managing multiple AI providers creates several headaches that slow down development and increase operational complexity.

The Multi-Provider Challenge

Each AI provider has its own SDK format, authentication method, and billing system. You end up maintaining separate code paths for each service, which makes testing new models tedious and switching providers a major refactoring project.

When providers experience downtime or rate limits, your application breaks, and there's no automatic failover. You're also flying blind on costs since each provider bills separately with different pricing structures.

The Observability Gap

Even after integrating multiple providers, you lack unified visibility. Request logs are scattered across different dashboards, making it impossible to compare model performance, track total costs, or understand usage patterns across your entire AI stack.

How Helicone AI Gateway Fixes This

Helicone AI Gateway addresses these problems through an integrated LLMOps approach:

These features work together to eliminate infrastructure complexity while giving you complete visibility into your AI operations.

Who Should Use Helicone AI Gateway?

Different teams get value from this unified approach. Here are some common use cases:

Prerequisites

Before diving into Helicone AI Gateway, you'll need a few things configured. This guide assumes you're comfortable with basic API development and have worked with REST APIs before.

Install the required packages:

# For Node.js/TypeScript
npm install openai dotenv

# For Python
pip install openai python-dotenv

Create a Helicone account at helicone.ai and generate an API key from the dashboard. This single API key gives you access to all providers!

Create a .env file in the root of your project with your Helicone API key:

HELICONE_API_KEY=sk-helicone-xxx

Making Your First API Call

To make your first LLM call you just need to change two things: the base URL and the API key.

Your First Request

import OpenAI from "openai";
import dotenv from "dotenv";

dotenv.config();

const client = new OpenAI({
  baseURL: "https://ai-gateway.helicone.ai",
  apiKey: process.env.HELICONE_API_KEY
});

const response = await client.chat.completions.create({
  model: "claude-4.5-haiku", // or 100+ other models - https://helicone.ai/models
  messages: [
    {
      role: "user",
      content: "Explain how Helicone AI Gateway works in one sentence"
    }
  ]
});

console.log(response.choices[0].message.content);

Understanding the Model Format

The model name follows a simple format that gives you flexibility in how requests are routed. Here are the common patterns:

// Automatic routing across all providers offering this model
model: "gpt-4o-mini"

// Route to a specific provider
model: "claude-sonnet-4/anthropic"

// Route to your custom deployment (for Azure, AWS, etc.)
model: "gpt-4o/azure/your-deployment-id"

// Route to specific fallback providers
model: "gpt-4o/openai,claude-sonnet-4/anthropic,gemini-2.5-flash/google"

You can use any model from the Helicone Model Registry.

What Makes This Different

Unlike other API gateways, Helicone automatically logs every request without additional configuration. Head to your Helicone Dashboard and you'll immediately see:

This built-in observability is what sets Helicone apart.

Intelligent Provider Routing

Building reliable AI applications means preparing for provider outages, rate limits, and unexpected failures.

Automatic Fallbacks

When you request a model without specifying a provider, the gateway tries all providers offering that model:

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    { role: "user", content: "What's the weather like today?" }
  ]
});

Manual Fallback Chains

For more control, you can specify a custom fallback sequence:

const response = await client.chat.completions.create({
  model: "gpt-4o/openai,claude-sonnet-4/anthropic,gemini-2.5-flash/google",
  messages: [
    { role: "user", content: "Analyze this business proposal..." }
  ]
});

Building Effective Fallback Strategies

Consider these patterns:

// Pattern 1: Cost-optimized with reliability
model: "gpt-4o-mini,claude-haiku-4,gemini-2.5-flash"

// Pattern 2: Capability-focused
model: "claude-opus-4-1,gpt-4o,gemini-2.5-pro"

// Pattern 3: Regional compliance
model: "claude-sonnet-4/azure/eu-deployment,gpt-4o/openai"

Monitoring Routing Decisions

Every request in your Helicone dashboard shows exactly which provider was used and why. You can see:

Working with Streaming Responses

When building user-facing AI features, especially for longer responses, users expect to see output appear progressively. Streaming solves this by sending response chunks:

Basic Streaming Setup

To enable streaming in Helicone AI Gateway, just add stream: true to your request:

const stream = await client.chat.completions.create({
  model: "claude-sonnet-4",
  messages: [
    { role: "user", content: "Write a detailed guide on prompt engineering" }
  ],
  stream: true
});

for await (const chunk of stream) {
  if (chunk.choices[0]?.delta?.content) {
    process.stdout.write(chunk.choices[0].delta.content);
  }
}

Building a Production Streaming Handler

async function streamResponse(model: string, messages: any[]) {
  const stream = await client.chat.completions.create({
    model,
    messages,
    stream: true
  });

let completeResponse = "";
  let tokenCount = 0;

for await (const chunk of stream) {
    const content = chunk.choices[0]?.delta?.content;
    if (content) {
      completeResponse += content;
      tokenCount++;
      process.stdout.write(content);
    }
  }

console.log(`\n\nStreaming complete: ${tokenCount} tokens`);
  return completeResponse;
}

Leveraging Observability Features

What truly sets Helicone AI Gateway apart is that observability is built into the core platform. Every request through the gateway is automatically logged, analyzed, and made queryable by default.

Automatic Request Logging

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [
    { role: "user", content: "Analyze this customer feedback..." }
  ]
});

Adding Custom Metadata

const client = new OpenAI({
  baseURL: "https://ai-gateway.helicone.ai",
  apiKey: process.env.HELICONE_API_KEY,
  defaultHeaders: {
    "Helicone-Session-Id": "chat-session-123",
    "Helicone-User-Id": "user-456"
  }
});

Session Tracking

const sessionId = crypto.randomUUID();
await client.chat.completions.create({
  model: "claude-sonnet-4",
  messages: [{ role: "user", content: "Hello!" }],
  extra_body: {
    helicone: {
      session: {
        id: sessionId
      }
    }
  }
});

Using Prompt Management

Helicone's prompt management system lets you update prompts without code changes or redeployments.

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  prompt_id: "sad98f45",
  inputs: {
    customer_name: customerName,
    issue_type: issueType
  }
});

Cost Optimization Strategies

Understanding Cost Visibility

Every request in your dashboard shows exact costs based on provider pricing. You can visualize costs by model, feature, user, and more.

Cost Optimization

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: messages
});

Monitoring Cost Trends

This visibility helps identify which models cost the most and where savings can be made.

Conclusion

We’ve covered how to access 100+ AI models through Helicone’s unified API Gateway, from making your first request to implementing advanced features like intelligent routing, streaming, observability, prompt management, and cost optimization.

Start with simple requests and gradually adopt features as your needs grow.