API documentation · v1

Hamro AI API

हाम्रो AI · API कागजात

One OpenAI-compatible API in front of many models. No model picking — auto-routing is always on, with instant failover if a model errors.

Get free API key

Free · Instant · No signup — POST /api/keys returns an og_ key in one request.

01 / authentication

Authentication

Hamro AI issues free og_ keys instantly — no account, no email, no credit card. Send the key as a standard Authorization: Bearer og_… header on any request. In the sandbox, anonymous calls are also allowed; in production a key is required and rate-limited.

Issue a key

bashPOST /api/keys
curl -X POST https://ai.hamro.site/api/keys \
  -H "Content-Type: application/json" \
  -d '{ "label": "my-app" }'

Response · 201

json
{
  "key": "og_9f3c2a1b8e7d6c5a4b3c2d1e0f9a8b7c6d5e4f3a2b1c",
  "prefix": "og_9f3c2a1…",
  "label": "my-app",
  "message": "Store this key securely — it is shown only once. Use it as: Authorization: Bearer og_..."
}

The raw key is shown once — store it securely. The backend only stores a SHA-256 hash and a display prefix.

Use it on a request

bash
curl https://ai.hamro.site/v1/chat/completions \
  -H "Authorization: Bearer og_9f3c2a1b8e7d6c5a4b3c2d1e0f9a8b7c6d5e4f3a2b1c" \
  -H "Content-Type: application/json" \
  -d '{ "messages": [ { "role": "user", "content": "नमस्ते!" } ] }'
02 / chat

Chat Completions

POST/v1/chat/completions

OpenAI-compatible. Send a messages[] array, optionally stream the response.

Request example

bash
curl https://ai.hamro.site/v1/chat/completions \
  -H "Authorization: Bearer og_..." \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      { "role": "system", "content": "You are a concise assistant." },
      { "role": "user", "content": "Write a haiku about Kathmandu mornings." }
    ],
    "temperature": 0.7,
    "max_tokens": 200
  }'

Response · non-stream · 200

json
{
  "id": "chatcmpl-j7t2nq9a4h",
  "object": "chat.completion",
  "created": 1735689600,
  "model": "Hamro Chat Mini",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Mist over Bagmati — / prayer bells stitch the dawn — / the city wakes soft." },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 0, "completion_tokens": 0, "total_tokens": 0 },
  "hamro": {
    "model_used": "Hamro Chat Mini",
    "provider_used": "hamro",
    "fallback_count": 0,
    "auto_routed": true
  }
}

Streaming · SSE

Send stream: true and the response comes back as text/event-stream — each token arrives as a data: { … } chunk, terminated by data: [DONE].

text
data: {"choices":[{"delta":{"role":"assistant","content":""}}]}

data: {"choices":[{"delta":{"content":"Mist"}}]}

data: {"choices":[{"delta":{"content":" over"}}]}

data: {"choices":[{"delta":{"content":" Bagmati"}}]}

...

data: [DONE]
03 / image

Image Generations

POST/v1/images/generations

Send a prompt and optional n, size. The response returns data[] with either a URL or b64_json.

Request

bash
curl https://ai.hamro.site/v1/images/generations \
  -H "Authorization: Bearer og_..." \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "A watercolor of Patan Durbar Square at dawn, soft pastels",
    "n": 2,
    "size": "1024x1024"
  }'

Response · 200

json
{
  "created": 1735689600,
  "data": [
    { "url": "https://cdn.hamro.site/img/9f3c2a1b.png" },
    { "b64_json": "iVBORw0KGgoAAAANSUhEUgAA…" }
  ],
  "hamro": {
    "model_used": "Hamro Image Pro",
    "provider_used": "hamro",
    "fallback_count": 0
  }
}
04 / unified

Smart Endpoint

POST/v1/smart

One endpoint — send { prompt } and the backend figures out the right kind of result.

Send a single prompt field. The router detects the task type from keyword rules and routes to chat, image, code, or audio. The response shape tells you which path was taken via task_type.

Task-type auto-detection

Prompt keyword patternDetected type
generate / draw / paint / render + image / picture / logo / banner …image
transcribe / audio / speech / tts / narrate / read aloudaudio
write code / function / class / component / script / algorithmcode
(anything else)chat

Request

bash
curl https://ai.hamro.site/v1/smart \
  -H "Authorization: Bearer og_..." \
  -H "Content-Type: application/json" \
  -d '{ "prompt": "draw an image of a rhododendron in watercolor" }'

Response · image task · 200

json
{
  "task_type": "image",
  "images": [
    { "url": "https://cdn.hamro.site/img/9f3c2a1b.png" }
  ],
  "model_used": "Hamro Image Pro",
  "provider_used": "hamro",
  "fallback_count": 0,
  "auto_routed": true
}

Response · chat task · 200

json
{
  "task_type": "chat",
  "content": "Namaste! How can I help you today?",
  "model_used": "Hamro Chat Mini",
  "provider_used": "hamro",
  "fallback_count": 0,
  "auto_routed": true
}

For chat-task prompts, stream: true returns the same SSE stream as /v1/chat/completions.

05 / models

Models

GET/v1/models

Lists enabled, online, free models — metadata only, no secrets. OpenAI-compatible shape.

bash
curl https://ai.hamro.site/v1/models

Response · 200 (truncated)

json
{
  "object": "list",
  "data": [
    {
      "id": "hamro-chat-mini",
      "object": "model",
      "created": 1735689600,
      "owned_by": "hamro",
      "hamro": {
        "display_name": "Hamro Chat Mini",
        "provider": "Hamro (built-in)",
        "task_type": "chat",
        "context_length": 32768,
        "health": "online",
        "auto_routed_only": true
      }
    },
    {
      "id": "hamro-image-pro",
      "object": "model",
      "created": 1735689600,
      "owned_by": "hamro",
      "hamro": {
        "display_name": "Hamro Image Pro",
        "provider": "Hamro (built-in)",
        "task_type": "image",
        "context_length": null,
        "health": "online",
        "auto_routed_only": true
      }
    }
    …
  ]
}
06 / status

Status

GET/v1/status

Machine-readable gateway status — summary, by-provider, models, and recent activity.

bash
curl https://ai.hamro.site/v1/status

Response · 200 (truncated)

json
{
  "generated_at": "2025-01-01T00:00:00.000Z",
  "summary": {
    "models_online": 4,
    "models_degraded": 0,
    "models_offline": 0,
    "total_models": 4,
    "total_requests": 1280,
    "total_fallback_events": 3,
    "estimated_value_saved_usd": 12.42,
    "success_rate": 0.9977
  },
  "by_provider": [
    { "slug": "hamro", "name": "Hamro (built-in)", "online": 4, "total": 4, "free": 4 }
    …
  ],
  "models": [
    { "id": "hamro-chat-mini", "displayName": "Hamro Chat Mini", "providerSlug": "hamro",
      "taskType": "chat", "status": "online", "avgLatencyMs": 412, "successRate": 1, "isFree": true, "enabled": true }
    …
  ],
  "recent": [ … ]
}
Model health updates from real traffic, and the gateway also runs scheduled active probes — so the status field is always fresh. What you see is what we route.
07 / usage

Usage

Per-key usage analytics. Send your og_ key and get back its real quota, totals, per-model split, a 14-day daily series and the 25 most recent requests — everything the /usage dashboard renders. Calling it never increments your daily counter.

GET/v1/usage

Your key's real usage — quota, success rate, savings, latency and recent requests.

Authentication required — send Authorization: Bearer og_…. A missing or unknown key returns 401.

bash
curl https://ai.hamro.site/v1/usage \
  -H "Authorization: Bearer og_YOUR_KEY"

Response · 200 (truncated)

json
{
  "generated_at": "2025-01-01T00:00:00.000Z",
  "key": {
    "prefix": "og_xxxxxx…", "label": "default",
    "created_at": "2025-01-01T00:00:00.000Z",
    "requests_today": 3, "daily_limit": 200, "daily_remaining": 197
  },
  "totals": {
    "requests": 12, "success": 11, "success_rate": 0.9167,
    "fallback_events": 2, "estimated_value_saved_usd": 0.04,
    "avg_latency_ms": 640
  },
  "by_task_type": { "chat": 9, "image": 2 },
  "by_model": [
    { "model": "Hamro GLM · Chat", "provider": "hamro", "requests": 8 }
    …
  ],
  "daily": [
    { "date": "2025-01-01", "requests": 4, "success": 4 }
    … 14 UTC days, zero-filled …
  ],
  "recent": [
    { "timestamp": "2025-01-01T00:00:00.000Z", "task_type": "chat",
      "requested_via": "playground", "model": "Hamro GLM · Chat",
      "provider": "hamro", "success": true, "latency_ms": 812,
      "fallback_count": 0 }
    …
  ]
}
08 / errors

Errors

All errors return JSON of the form { "error": { "message": "…" } } with the appropriate HTTP status. The 503 message is friendly and safe to surface to end-users.

StatusMeaningExample message
400Bad request bodyinvalid JSON body · messages[] or prompt required
401Invalid or unknown API keyinvalid key format · unknown key
429Daily limit reacheddaily limit reached · see Retry-After
500Internal server errorinternal error
503All providers exhaustedAll providers exhausted. Please retry in a moment.
09 / recipes

Quickstart Recipes

Complete, copy-paste recipes for real API consumers — the official openai SDK in Python and Node, plus a streaming cURL. Every snippet runs as-is; grab a free og_ key at /get-key first.

Python · openai SDK

The official openai package works unmodified — point base_url at the gateway and drop your og_ key in. The model field just keeps the SDK happy; the router always picks.

pythonquickstart.py
from openai import OpenAI

client = OpenAI(
    base_url="https://ai.hamro.site/v1",
    api_key="og_YOUR_KEY",  # free at /get-key
)

response = client.chat.completions.create(
    model="auto",  # ignored — the router picks
    messages=[{"role": "user", "content": "Namaste! Introduce yourself."}],
)
print(response.choices[0].message.content)

Node.js · openai SDK

Same story in Node — set baseURL, read the key from an environment variable, and await the completion. Top-level await works out of the box in an .mjs file.

javascriptquickstart.mjs
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://ai.hamro.site/v1",
  apiKey: process.env.HAMRO_KEY, // free at /get-key
});

const response = await client.chat.completions.create({
  model: "auto", // ignored — the router picks
  messages: [{ role: "user", content: "Namaste! Introduce yourself." }],
});
console.log(response.choices[0].message.content);

cURL · streaming

No SDK required — -N disables curl's buffering so each SSE token prints to your terminal the moment it arrives.

bashstreaming
curl -N https://ai.hamro.site/v1/chat/completions \
  -H "Authorization: Bearer og_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "stream": true,
    "messages": [{"role": "user", "content": "Write a haiku about the Himalayas"}]
  }'
10 / limits

Rate Limits

Free keys carry a per-key limit and a per-IP limit. Both are token-bucket enforced. When a 429 fires, the response includes a Retry-After header in seconds.

Per key
200 / day
30 / min
Per IP
200 / day
30 / min

On 429 the response carries a Retry-After header (in seconds) — wait that long before retrying.

Daily quota headers

Every response made with a known key — success or error, on every /v1 endpoint — carries three x-ratelimit-* headers computed live from your key's record (the same numbers as GET /v1/usage). The remaining count already includes the request that produced the response; the quota resets at 00:00 UTC.

HeaderMeaning
x-ratelimit-limit-dayYour key's daily request limit (200 on a fresh key).
x-ratelimit-remaining-dayRequests remaining today — limit minus today's UTC-day count, floored at 0.
x-ratelimit-reset-dayISO-8601 timestamp of the next UTC-midnight reset — when remaining refills to the limit.
retry-after429 responses only — seconds until that reset. Wait, then retry.
bashuse -i to see headers
curl -si -X POST https://ai.hamro.site/v1/chat/completions \
  -H "authorization: Bearer og_YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"messages":[{"role":"user","content":"say hi"}]}'

Example response headers

text
# 200 OK — on every authenticated response
x-ratelimit-limit-day: 200
x-ratelimit-remaining-day: 197
x-ratelimit-reset-day: 2025-01-02T00:00:00.000Z

# 429 Too Many Requests — daily limit reached, adds retry-after
retry-after: 43180
x-ratelimit-limit-day: 200
x-ratelimit-remaining-day: 0
x-ratelimit-reset-day: 2025-01-02T00:00:00.000Z
11 / headers

Auto-routing Headers

Every successful response — chat, image, and smart — includes three x-gateway-* headers so you always know who really served you. The same fields are echoed in the hamro object of the JSON body.

HeaderMeaning
x-gateway-model-usedDisplay name of the model that actually served the request.
x-gateway-provider-usedSlug of the provider that served the request.
x-gateway-fallback-countNumber of higher-priority models that failed before this one answered. 0 means the primary answered. Greater than 0 means the primary failed and a backup answered in milliseconds.
textexample response headers
HTTP/1.1 200 OK
content-type: application/json
x-gateway-model-used: Hamro Chat Mini
x-gateway-provider-used: hamro
x-gateway-fallback-count: 0
12 / sdks

SDKs

The OpenAI SDK works unmodified — just point it at https://ai.hamro.site/v1 and pass your og_ key as the API key. Drop-in replacement.

JavaScript / TypeScript

javascript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://ai.hamro.site/v1",
  apiKey: "og_9f3c2a1b8e7d6c5a4b3c2d1e0f9a8b7c6d5e4f3a2b1c",
});

const res = await client.chat.completions.create({
  // model is required by the SDK type but ignored by the gateway.
  model: "auto",
  messages: [{ role: "user", content: "नमस्ते! के खबर?" }],
});

console.log(res.choices[0].message.content);

Python

python
from openai import OpenAI

client = OpenAI(
    base_url="https://ai.hamro.site/v1",
    api_key="og_9f3c2a1b8e7d6c5a4b3c2d1e0f9a8b7c6d5e4f3a2b1c",
)

res = client.chat.completions.create(
    # model is required by the SDK but ignored by the gateway.
    model="auto",
    messages=[{"role": "user", "content": "नमस्ते! के खबर?"}],
)

print(res.choices[0].message.content)
Try it now

Send your first request.

Try streaming chat or image generation in the playground — no key required in the sandbox.