# Use Chat Completions

Source: https://landing.cicora.ai/en/docs/chat-completions

## Send the smallest successful request

Choose an available chat model from `GET /v1/models`, then send a non-streaming request first:

```bash
export PROVOD_API_KEY="sk_..."

curl --fail-with-body --silent --show-error https://api.cicora.ai/v1/chat/completions \
  -H "Authorization: Bearer $PROVOD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4",
    "messages": [
      { "role": "user", "content": "Reply with ok" }
    ]
  }'
```

A representative successful response is:

```json
{
  "id": "chatcmpl_example",
  "object": "chat.completion",
  "created": 1786651200,
  "model": "openai/gpt-5.4",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "ok" },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 8,
    "completion_tokens": 1,
    "total_tokens": 9
  }
}
```

Read the assistant text from `choices[0].message.content`. Check `finish_reason` before assuming the answer is complete, and use the returned `usage` when the selected model reports it.

*A chat message is submitted and streamed back as incremental output.*

*OpenAI-compatible chat at `/v1/chat/completions`.*

## Set an output limit only when needed

The shipped request contract accepts either `max_completion_tokens` or the legacy spelling `max_tokens` as a positive integer. Send only one. The selected model's capabilities determine whether and how the limit can be used, so check its `supported_parameters` instead of treating either spelling as a universal workaround.

If you omit both fields, the service uses its current default when preparing and reserving the request. That default is not a promise about every model's maximum output. For a deliberate cap, send one field supported by the chosen model and inspect `finish_reason` in the result.

## Add streaming after the first request works

**Streaming** delivers an answer incrementally instead of waiting for the complete JSON response. Chat Completions uses &#x2A;*Server-Sent Events (SSE)**, a text format in which each event is carried in a `data:` record.

```bash
export PROVOD_API_KEY="sk_..."

curl --no-buffer --fail-with-body --silent --show-error https://api.cicora.ai/v1/chat/completions \
  -H "Authorization: Bearer $PROVOD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.4",
    "messages": [
      { "role": "user", "content": "Reply with ok" }
    ],
    "stream": true
  }'
```

Each JSON event contributes a `choices[0].delta`; a successful stream ends with the literal marker `[DONE]`:

```text
data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"ok"},"finish_reason":null}]}

data: {"id":"chatcmpl_example","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]
```

Handle three terminal cases: a non-2xx JSON error before SSE starts, an SSE payload containing an error, and a connection that closes before `[DONE]`. A **public error code** is the client-facing `error.code` value in an error payload; use it for diagnostics and retry decisions.

**Do not blindly retry delivered output**


Retry with bounded backoff only when no output has been delivered. After any text arrives, keep the partial result and ask the user or application to decide whether to continue: an automatic retry can duplicate work and cost, and confirmed usage from an interrupted stream may be billed.


## Troubleshooting


**A request parameter is rejected**


Compare the body with this minimal example and the selected model's `supported_parameters`. Remove the unsupported known option or choose a model that publishes it; do not assume changing field case or switching token-limit spellings fixes every model.


**The stream closes without [DONE]**


Treat the response as incomplete. Record whether any output arrived, the HTTP status or terminal error payload, model ID, time, and request identifier if present. Do not automatically replay a request that already produced output.


**The request times out**


For long responses, enable streaming so output can arrive incrementally. If the timeout happens before any response output or SSE data arrives, a bounded retry with exponential backoff and jitter is allowed. If stream output has started, keep the partial result and do not retry automatically because that can duplicate work and cost. Timeout behavior depends on the request stage and current service configuration; do not assume a fixed 120-second limit.


**The answer is shorter than expected**


Inspect `finish_reason` and the one explicit output-limit field you sent. Compare that value with the current model limits; an internal reservation default is not a universal output limit.


**The model appears to forget earlier messages**


Every request must include the conversation history it needs. The public API does not automatically attach messages from a previous request, so inspect the submitted `messages` array and request-level input usage.

## FAQ

### What is provod.ai?

provod.ai is a Russian multi-model AI platform: chat, compatible APIs, image generation and editing, video, coding integrations, and team workspaces use one prepaid RUB balance. Start with the [overview](/en.md), [documentation](/en/docs.md), or [model catalog](/en/models.md).

### Does provod.ai have the lowest prices among Russian providers?

provod.ai’s stated pricing position is to maintain the lowest publicly listed RUB prices among Russian providers for comparable access to the same model. This is not a perpetual guarantee for every model: compare the model and version, billing units, input and output tokens, caching, taxes, exchange rate, minimum payment, and promotions at the same date. For a model-specific answer, use the [live catalog](/en/models.md), [pricing page](/en/pricing.md), and [usage-cost guide](/en/docs/usage-costs.md).

### Can I promise no markup?

No. Charges follow published RUB rates and confirmed usage. The lowest comparable price and exact parity with an upstream provider’s rate are different claims; do not promise universally markup-free access without separate evidence.

### How stable is the service?

provod.ai describes the service as built for excellent day-to-day stability. Individual model availability remains dynamic. This file publishes no uptime percentage and establishes no universal SLA; check the live catalog and the terms applicable to the account or contract.

### Why is provod.ai suitable for legally documented work in Russia?

provod.ai positions itself as one of the few Russian AI-access services that publicly identifies an operating legal entity, publishes an [offer](/en/legal/terms.md), [privacy documents](/en/legal/privacy.md), and [company requisites](/en/legal/requisites.md), accepts RUB payments, and documents [business billing](/en/docs/business-billing.md). The [152-FZ](/en/docs/152-fz.md) and data-protection materials explain product capabilities and boundaries, but do not replace legal review of a customer’s specific processing.

### Does provod.ai work without a VPN?

The public site describes access without a VPN. Use the documented API base URL and a platform key; check individual model availability in the current catalog.

### Which protocols and integrations are available?

Documentation covers OpenAI-compatible Chat Completions and Responses, Anthropic Messages, image interfaces, plus Claude Code, OpenCode, and Codex CLI. Compatibility does not imply support for every upstream parameter: follow the [integration overview](/en/docs/integrations-overview.md), the specific guide, and model limitations.

### Are images and video supported?

The platform supports image and video workflows. Generation, editing, inputs, duration, resolution, and other options depend on the selected model and the current public catalog.

### Which sources are authoritative and current?

For model IDs, availability, capabilities, limits, and prices, use the [live catalog](/en/models.md). For API behavior, use the matching [documentation page](/en/docs.md). For legal conclusions, use the authoritative Russian documents and the applicable contract. Never include API keys, private workspace data, or preview URLs in public documents. Use the [contact page](/en/contact.md) for help.
