OpenAI API in Code: Chat Completions, Responses, Tool Calls, Structured Outputs and Retries
Key takeaways
How to call the OpenAI API from Python and Node with the v1+ SDK client: Chat Completions and the Responses API, tools with tool_choice instead of the deprecated functions parameter, Structured Outputs, streaming, and timeouts and retries that do not multiply your bill.
What this post assumes
This is a working reference for calling the OpenAI API from backend code with the official openai SDKs (Python and Node, v1 and later). I checked every method and parameter name below against the type definitions of the installed packages (openai 3.19.2 for Python and 7.23.0 for Node at the time of writing). A lot of tutorials still floating around were written for the pre-1.0 Python SDK (openai.ChatCompletion.create, openai.api_key = ..., request_timeout=), and that code does not run on a current install.
What this post deliberately does not do is quote prices, rate-limit tiers, context sizes, or recommend a specific model. Those change more often than blog posts get updated, and a stale price table is worse than none. Check the models page and the pricing page when you pick a model, and the limits page in your organization’s dashboard for your actual rate limits. In the code, the model id comes from an environment variable:
export OPENAI_API_KEY="sk-..." # never commit this
export OPENAI_MODEL="<a model id from the models page>"
pip install openai # Python
npm install openai # Node
Creating the client once
Both SDKs read OPENAI_API_KEY from the environment if you do not pass api_key, so you rarely need to touch the key in code. Create one client at startup and reuse it; it holds a connection pool.
import os
from openai import OpenAI
client = OpenAI(
timeout=30.0, # seconds; the SDK default is 600s, far too long for a web request
max_retries=2, # this is the default; set it explicitly so it is visible in review
)
MODEL = os.environ["OPENAI_MODEL"]
import OpenAI from 'openai';
const client = new OpenAI({
timeout: 30_000, // milliseconds in the Node SDK
maxRetries: 2,
});
const MODEL = process.env.OPENAI_MODEL;
Note the unit difference: Python takes seconds (or an httpx timeout object for separate connect/read values), Node takes milliseconds. The old request_timeout argument is gone. For a single slow call, override per request instead of loosening the whole client: client.with_options(timeout=120).chat.completions.create(...) in Python, or pass { timeout: 120_000 } as the second argument in Node.
Chat Completions: the stateless basics
Chat Completions takes a list of messages and returns choices. The API keeps no conversation state between calls; the model only sees what is in messages for this request.
completion = client.chat.completions.create(
model=MODEL,
messages=[
{"role": "system", "content": "You answer questions about our billing API. Be brief."},
{"role": "user", "content": "How do I rotate a webhook secret?"},
],
max_completion_tokens=500,
)
choice = completion.choices[0]
print(choice.message.content)
print(choice.finish_reason) # "stop", "length", "tool_calls", "content_filter"
print(completion.usage.total_tokens) # log this; it is what you pay for
Two details that matter in production:
max_tokensis marked deprecated in the SDK in favor ofmax_completion_tokens, and it is not accepted by reasoning models. Use the new name in new code.- Check
finish_reason."length"means the answer was cut off at your token cap. If you are about tojson.loadsthat content, it will fail, and the right fix is usually a bigger cap or a shorter requested output, not a retry.
Because the API is stateless, a chat UI must resend history every turn. That history is billed as input tokens on every request, so a long conversation gets more expensive per turn as it grows. Keep the system message, trim or summarize old turns, and cap by token count rather than message count.
def trim_history(messages, keep_last=12):
system = [m for m in messages if m["role"] == "system"]
rest = [m for m in messages if m["role"] != "system"]
return system + rest[-keep_last:]
This naive version can split a tool call from its tool result, which the API rejects. If you use tools, trim at user-turn boundaries.
The Responses API
The Responses API (client.responses.create) is OpenAI’s newer interface. It takes instructions plus input (a string or a list of input items), returns a list of typed output items, and exposes an output_text convenience property that joins the text parts:
response = client.responses.create(
model=MODEL,
instructions="You answer questions about our billing API. Be brief.",
input="How do I rotate a webhook secret?",
max_output_tokens=500,
)
print(response.output_text)
const response = await client.responses.create({
model: MODEL,
instructions: 'You answer questions about our billing API. Be brief.',
input: 'How do I rotate a webhook secret?',
});
console.log(response.output_text);
It can also carry conversation state server-side via previous_response_id (or the conversation parameter), so you do not have to resend history yourself. Whether you want that depends on your data-retention requirements; if you already store conversations in your own database, Chat Completions with explicit history is simpler to reason about. Both APIs are supported by the SDK. The rest of this post shows Chat Completions first because it is what most existing code uses, and notes the Responses equivalent where the shape differs.
Tool calling with tools and tool_choice
The old functions / function_call parameters and the role: "function" message are deprecated. The current shape:
- You describe tools in
tools, each wrapped as{"type": "function", "function": {...}}. - The model may reply with
message.tool_calls, a list. It can ask for several calls in one turn. - You run each call and append a
role: "tool"message with the matchingtool_call_id. - You call the model again with the extended history.
import json
tools = [
{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the shipping status of an order by its id.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
"additionalProperties": False,
},
"strict": True,
},
}
]
def get_order_status(order_id: str) -> dict:
return {"order_id": order_id, "status": "shipped"} # your real lookup here
HANDLERS = {"get_order_status": get_order_status}
def run(user_text: str, max_rounds: int = 5) -> str:
messages = [{"role": "user", "content": user_text}]
for _ in range(max_rounds):
completion = client.chat.completions.create(
model=MODEL, messages=messages, tools=tools, tool_choice="auto",
)
msg = completion.choices[0].message
if not msg.tool_calls:
return msg.content
messages.append(msg) # the assistant turn that contains the tool calls
for call in msg.tool_calls:
handler = HANDLERS.get(call.function.name)
try:
args = json.loads(call.function.arguments)
result = handler(**args) if handler else {"error": "unknown tool"}
except Exception as exc: # bad JSON or bad args: tell the model, do not crash
result = {"error": str(exc)}
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
raise RuntimeError("tool loop did not finish")
Things that are easy to get wrong:
- Every tool call needs an answer. If the model returns three calls and you append only one tool message, the next request fails validation. Loop over all of them.
strict: Trueasks the model to follow the JSON Schema exactly. Strict mode needsadditionalProperties: falseand every property listed inrequired(use a nullable type for optional fields). Without it, treatargumentsas untrusted JSON that may be missing fields.tool_choicecan be"auto"(default when tools are present),"none","required", or{"type": "function", "function": {"name": "get_order_status"}}to force one tool.- Cap the loop. A model that keeps calling tools will keep costing money.
max_roundsis a cheap guard. - Tool arguments are model output. Validate them like user input. A tool that runs SQL or shell commands built from
argumentsis a prompt-injection hole.
In Python, openai.pydantic_function_tool(MyModel) builds the tool definition from a Pydantic model; the Node SDK has zodFunction in openai/helpers/zod.
In the Responses API the shape is flatter: a tool is {"type": "function", "name": ..., "parameters": ..., "strict": ...}, the model emits output items of type function_call with call_id, name and arguments, and you send results back as input items {"type": "function_call_output", "call_id": ..., "output": "..."}.
Structured Outputs instead of “please reply in JSON”
Asking for JSON in the prompt works most of the time, which is exactly what makes it dangerous. The SDKs now offer two stronger options:
- JSON mode:
response_format={"type": "json_object"}. The SDK docs call this the older mode. It guarantees valid JSON syntax, not the fields you need, and you still have to say “JSON” in your instructions. - Structured Outputs:
response_format={"type": "json_schema", "json_schema": {...}}, which makes the model match your schema. The SDK docs call this the preferred option on models that support it.
The parse helpers wrap Structured Outputs around a Pydantic or Zod model:
from pydantic import BaseModel
class Ticket(BaseModel):
category: str
priority: int
summary: str
completion = client.chat.completions.parse(
model=MODEL,
messages=[
{"role": "system", "content": "Classify the support email."},
{"role": "user", "content": email_body},
],
response_format=Ticket,
)
msg = completion.choices[0].message
if msg.refusal:
handle_refusal(msg.refusal)
else:
ticket: Ticket = msg.parsed
import { z } from 'zod';
import { zodResponseFormat } from 'openai/helpers/zod';
const Ticket = z.object({
category: z.string(),
priority: z.number().int(),
summary: z.string(),
});
const completion = await client.chat.completions.parse({
model: MODEL,
messages: [
{ role: 'system', content: 'Classify the support email.' },
{ role: 'user', content: emailBody },
],
response_format: zodResponseFormat(Ticket, 'ticket'),
});
const ticket = completion.choices[0].message.parsed;
The Responses API equivalent is client.responses.parse(..., text_format=Ticket) in Python and zodTextFormat in Node.
Even with a schema, keep two checks: refusal (the model declined) and truncation. If the output hits the token cap, the Python helper raises LengthFinishReasonError rather than returning half an object. A schema guarantees shape, not truth: a priority of 5 is valid JSON even if the email was a thank-you note.
Streaming
Streaming improves perceived latency, not total time or cost. Set stream=True and iterate:
stream = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Explain idempotency keys in two paragraphs."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.choices: # the final usage chunk has an empty choices list
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
if chunk.usage:
print("\n", chunk.usage.total_tokens, "tokens")
const stream = await client.chat.completions.create({
model: MODEL,
messages: [{ role: 'user', content: 'Explain idempotency keys in two paragraphs.' }],
stream: true,
stream_options: { include_usage: true },
});
for await (const chunk of stream) {
const delta = chunk.choices[0]?.delta?.content;
if (delta) process.stdout.write(delta);
if (chunk.usage) console.log('\n', chunk.usage.total_tokens, 'tokens');
}
With the Responses API you pass stream=True and switch on event types: text arrives in response.output_text.delta events and the final object in response.completed.
I have broken streaming in the same few ways more than once, and they are all well-known. The first is the include_usage chunk: turning on usage reporting adds a final chunk whose choices array is empty, so the innocent-looking chunk.choices[0] in every tutorial raises an IndexError on the very last chunk, after the user has already seen the whole answer. The second is tool calls: when the model streams a tool call, the arguments string arrives in fragments keyed by index, and calling json.loads on each fragment fails. You have to concatenate the fragments per index and parse only after the stream ends. The third is proxying to a browser as Server-Sent Events: a delta that contains a newline breaks the data: framing unless you JSON-encode each piece, and a reverse proxy that buffers responses makes the whole “stream” arrive at once. Also note from the SDK docs: if the stream is interrupted, you may never receive the usage chunk, so do not rely on it alone for billing.
A minimal FastAPI proxy that avoids the framing problem:
import json
from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from openai import AsyncOpenAI
app = FastAPI()
aclient = AsyncOpenAI(timeout=60.0)
@app.post("/api/chat/stream")
async def chat_stream(body: dict):
async def events():
stream = await aclient.chat.completions.create(
model=MODEL, messages=body["messages"], stream=True,
)
async for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
yield f"data: {json.dumps({'delta': chunk.choices[0].delta.content})}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(events(), media_type="text/event-stream")
Use AsyncOpenAI inside async frameworks. Calling the sync client from an async def handler blocks the event loop for the whole generation.
Errors, 429s and retries
The SDK raises typed exceptions. The ones worth handling separately:
| Exception (Python) | Meaning | Retry? |
|---|---|---|
APITimeoutError | Your client timeout fired | SDK already retried |
APIConnectionError | Network problem | SDK already retried |
RateLimitError (429) | Rate limit hit, or quota exhausted | Depends on the error code |
BadRequestError (400) | Invalid request (schema, too many tokens, bad tool history) | No, fix the request |
AuthenticationError (401) | Bad or revoked key | No |
InternalServerError (5xx) | Server side | SDK already retried |
The built-in behavior, from the SDK source: by default the client retries twice on connection errors, 408, 409, 429 and 5xx, with exponential backoff and jitter, and it honors retry-after / retry-after-ms headers when the server sends them. Node’s RateLimitError and friends are exported on the OpenAI namespace with the same meaning.
So the most common mistake is writing a retry loop around a client that already retries:
from openai import RateLimitError, APIStatusError
try:
completion = client.chat.completions.create(model=MODEL, messages=messages)
except RateLimitError as e:
if e.code == "insufficient_quota":
alert_billing() # retrying will never fix this
else:
enqueue_for_later(messages) # SDK retries were exhausted; back off at the job level
except APIStatusError as e:
log.error("openai error", status=e.status_code, request_id=e.request_id)
raise
The cost version of that mistake is the one I would warn about first. With max_retries=2 on the client, three attempts in your own loop, and a job queue that retries failed jobs three more times, one logical request can turn into dozens of HTTP calls. A timeout that fires after the server has already generated most of the answer still gets billed, so a too-short timeout combined with stacked retries pays for the same answer several times and never delivers it. Pick one layer to own retries (usually the SDK for transient errors, your queue for “try again in a minute”), log the request_id on every failure so you can match it with support, and track usage from successful responses so a retry storm shows up on a dashboard before it shows up on the invoice.
A 429 also comes in two kinds. One is a rate limit: slow down and it will pass. The other is insufficient_quota: your account is out of credit or over its spend limit, and no amount of backoff fixes it. Check the error code before deciding.
Keep the key on the server
The API key is a bearer credential tied to your billing. It must never ship in browser or mobile app code, and NEXT_PUBLIC_ or VITE_ prefixed environment variables are compiled into the client bundle. The Node SDK refuses to run in a browser unless you pass dangerouslyAllowBrowser: true, and the name is the warning.
The pattern is a thin backend endpoint that holds the key, authenticates your own user, enforces per-user limits, and forwards a request built on the server:
// app/api/chat/route.ts (Next.js route handler, runs on the server)
import OpenAI from 'openai';
const client = new OpenAI({ timeout: 30_000 });
export async function POST(req) {
const user = await requireSession(req); // your auth
await enforceQuota(user.id); // your per-user limit
const { question } = await req.json();
const completion = await client.chat.completions.create({
model: process.env.OPENAI_MODEL,
messages: [
{ role: 'system', content: 'You answer questions about our product.' },
{ role: 'user', content: String(question).slice(0, 4000) },
],
max_completion_tokens: 600,
});
return Response.json({ answer: completion.choices[0].message.content });
}
Do not let the client send arbitrary messages including the system prompt, or choose the model and token cap. Otherwise your endpoint is a free, unauthenticated OpenAI proxy on your bill. Keys that do leak tend to leak through a public repo, a frontend bundle, or a log line that printed request headers, so also scrub Authorization from your logging. If a key does get out, revoke it in the dashboard first and investigate second; usage from a leaked key keeps accruing while you read logs.
Migrating old snippets
| Old (pre-1.0 or deprecated) | Current |
|---|---|
openai.api_key = "..." then openai.ChatCompletion.create(...) | client = OpenAI() then client.chat.completions.create(...) |
request_timeout=30 | OpenAI(timeout=30) or client.with_options(timeout=30) |
functions=[...], function_call="auto" | tools=[{"type": "function", "function": {...}}], tool_choice="auto" |
message.function_call | message.tool_calls (a list) |
{"role": "function", "name": ..., "content": ...} | {"role": "tool", "tool_call_id": ..., "content": ...} |
max_tokens | max_completion_tokens (Chat), max_output_tokens (Responses) |
| “Reply only in JSON” in the prompt | response_format with json_schema, or .parse(...) |
openai.error.RateLimitError | openai.RateLimitError |
Related Articles
- Claude Messages API: max_tokens, streaming, tool use loops, caching and retries covers the same production concerns for Anthropic’s API, useful if you support both providers.
- LangChain 1.x in practice covers when a framework on top of these SDKs helps and when it only adds layers.
- Prompt engineering for developers goes deeper on writing the system and user messages themselves.
Official references: OpenAI API docs, openai-python, openai-node.