The hosted inference API for large language models
Low-latency completions, function calling, and realtime token streaming over a single Bearer-authenticated endpoint. Ship to production without running your own GPUs.
Realtime streaming console
A live demo session streams model output token-by-token over a WebSocket. The public demo is read-only; your own sessions authenticate with an API key.
opening realtime session…
Quickstart
All API calls require a Bearer API key. Create one in the dashboard, then:
curl https://llm.api.jacewicz.org/v1/chat/completions \
-H "Authorization: Bearer $NORTHWIND_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "northwind-large",
"stream": true,
"messages": [{"role": "user", "content": "Say hello."}]
}'
Requests without a valid key return 401. Streaming responses are delivered as
server-sent chunks; realtime bidirectional sessions use the WebSocket console endpoint.
Models
northwind-large 128k context
northwind-turbo 32k context
northwind-embed embeddings