Skip to main content
For agents built on the higher-level OpenAI Agents SDK instead of the plain client, see OpenAI Agents SDK. For an OpenAI-compatible endpoint served by NVIDIA NIM, use patch_nim_client instead - same machinery, but traces are attributed to nvidia-nim rather than openai. Install the integration extra:

Usage

Call patch_openai_client() once after creating your OpenAI client. All subsequent client.chat.completions.create() calls are traced automatically. No changes to individual API calls are needed. Works with both openai.OpenAI and openai.AsyncOpenAI.
Async client:

What gets traced

By default, each chat.completions.create() call produces its own trace.

Streaming

Calls made with stream=True are traced too. The patch returns a transparent proxy over the provider’s stream: iterating it, using it as a context manager, calling close(), and reading its attributes all pass straight through to the real stream, and the trace is assembled from the chunks your code actually consumes.
What a streamed trace carries: If your code stops reading early (break) or an error interrupts the stream, the trace records what was streamed up to that point, and an error is recorded as the trace’s error. A stream that is dropped without being closed still records what it saw when it is garbage-collected. The raw OpenAI SDK has no built-in concept of “tool call” or “retrieval”; those only exist as plain Python code around your chat.completions.create() calls, so the patch can’t see them on its own. Use tracer.trace_tool_call() and tracer.trace_retrieval() to record them manually so they show up in the trace’s performance summary - see the Anthropic integration’s tool-use example for the same pattern (identical API, different client).

Multi-call agentic loops

Like the Anthropic integration, wrap a multi-call tool-use loop in with tracer.trace(...) to collapse every chat.completions.create() call made inside it into one trace instead of one trace per call:
This works because patch_openai_client() checks tracer.current_span on every call: if a span is active in the current context (a ContextVar, inherited by asyncio tasks), the call is attached to it as an "LLM Call N" step; otherwise it sends its own trace as usual. A patched call made outside any active span opens its own root trace, stamped with span kind llm - it is a bare model call.

patch_openai_client() reference

Calling patch_openai_client() on an already-patched client is a no-op; it is safe to call multiple times.
Call tracer.flush() before your process exits in scripts or one-shot jobs. In long-running servers it is not required: traces drain automatically in the background.