The hosted inference API for large language models

Low-latency completions, function calling, and realtime token streaming over a single Bearer-authenticated endpoint. Ship to production without running your own GPUs.

Realtime streaming console

A live demo session streams model output token-by-token over a WebSocket. The public demo is read-only; your own sessions authenticate with an API key.

opening realtime session…

Quickstart

All API calls require a Bearer API key. Create one in the dashboard, then:

curl https://llm.api.jacewicz.org/v1/chat/completions \
  -H "Authorization: Bearer $NORTHWIND_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "northwind-large",
    "stream": true,
    "messages": [{"role": "user", "content": "Say hello."}]
  }'

Requests without a valid key return 401. Streaming responses are delivered as server-sent chunks; realtime bidirectional sessions use the WebSocket console endpoint.

Models

northwind-large 128k context
northwind-turbo 32k context
northwind-embed embeddings