SuperPenguin Docs
SDKsTypeScript

Agent SDKs and dynamic metadata

Attribute Vercel, OpenAI Agents, LangGraph, Pi, OpenCode and custom harness model calls to each customer and run.

These APIs are implemented for the next @superpenguin/js release, version 0.11.0. They are not available in 0.9.1 or 0.10.0. Install 0.11.0 after it is published, or build the development branch.

SuperPenguin records each model invocation and groups it using the customer and agent-run metadata you supply. Calls still go directly to the model provider. Instrument every model or provider client that the agent can use, including those called inside tools and subagents.

Static and dynamic metadata

Static metadata belongs on the wrapped client, middleware, callback or Pi stream wrapper. Dynamic metadata belongs around the in-process agent invocation. OpenCode uses a server-side session resolver, described below:

import { init, withMetadata } from "@superpenguin/js";

init({ apiKey: process.env.SP_API_KEY });

await withMetadata(
  {
    customer_id: customer.id,
    custom_tags: { ticket_id: ticket.id, agent_run_id: runId },
  },
  () => agent.generate({ prompt }),
);

withMetadata uses Node.js execution-local storage. Concurrent runs can share an agent or model without sharing customer metadata. Nested scopes merge custom tags. Metadata precedence is static defaults, then nested execution scopes, then explicit per-call metadata. Native wrappers accept spMetadata; LangChain callbacks accept invocation metadata. currentMetadata() returns a detached snapshot for a custom integration.

The metadata is copied when the model invocation starts, so delayed streaming completions retain the right customer. Keep agent stream creation and consumption inside the scope, since consumption can start additional model calls:

await withMetadata({ customer_id: customer.id }, async () => {
  const result = await agent.stream({ prompt });
  for await (const text of result.textStream) {
    sendToUser(text);
  }
});

This is a Node.js server API. It does not propagate across a queue, worker thread, process, remote agent, or a paused run resumed in a new execution. Include attribution metadata in the job or resume payload, then establish a new scope at the destination. Framework context, trace metadata and conversation history are not automatically copied into SuperPenguin metadata.

Vercel AI SDK

Wrap each language model with superpenguinMiddleware. It intercepts every model call made by the agent, including calls after a tool result. For AI SDK 7, select middleware version v4; AI SDK 6 uses v3 (the default).

import { openai } from "@ai-sdk/openai";
import { ToolLoopAgent, wrapLanguageModel } from "ai";
import { superpenguinMiddleware, withMetadata } from "@superpenguin/js";

const model = wrapLanguageModel({
  model: openai("gpt-4o-mini"),
  middleware: superpenguinMiddleware({
    specificationVersion: "v4", // AI SDK 7; use "v3" for AI SDK 6.
    provider: "openai",
    metadata: { feature: "support_agent", environment: "production" },
  }),
});
const agent = new ToolLoopAgent({ model, instructions: "Answer support questions." });

const result = await withMetadata(
  { customer_id: "acme", custom_tags: { ticket_id: "123" } },
  () => agent.generate({ prompt: "How do I reset my password?" }),
);

Usage is captured per model call. Gateway-reported costs are forwarded when available; otherwise SuperPenguin prices captured tokens. Streaming middleware forwards original chunks with backpressure and captures usage from the consumed finish chunk. The finish chunk is delivered before telemetry settles; stream closure awaits telemetry to keep request-scoped runtimes alive, and telemetry failures preserve the original stream. If a stream is cancelled before usage arrives, no completed zero-token event is invented. Unreported provider usage cannot be recovered from a cancelled stream.

Use one capture path per invocation. Do not combine model middleware with trackGenerateText, trackStreamText, wrapped underlying clients, or an exporter that reports the same call. The text helpers also capture individual result steps when available, but the middleware is preferable for agents, changing models and calls inside tools. All models selected by prepareStep must be instrumented.

References: model middleware and agent configuration.

OpenAI Agents SDK

Inject a wrapped OpenAI client into a runner's model provider. This observes the actual Responses or Chat Completions calls; agent tracing is not required. OpenAI Agents 0.19 uses OpenAI SDK 7, so use compatible package versions.

import OpenAI from "openai";
import { Agent, OpenAIProvider, Runner } from "@openai/agents";
import { wrap, withMetadata } from "@superpenguin/js";

const client = wrap(new OpenAI(), {
  metadata: { feature: "support_agent", environment: "production" },
});
const runner = new Runner({
  modelProvider: new OpenAIProvider({ openAIClient: client }),
});
const agent = new Agent({ name: "Support", model: "gpt-4o-mini" });

const result = await withMetadata(
  { customer_id: "acme", custom_tags: { ticket_id: "123" } },
  () => runner.run(agent, "How do I reset my password?"),
);

The runner's context remains your application context. Select the scalar fields to attribute and pass them to withMetadata; SuperPenguin does not serialize the entire context. traceMetadata does not substitute for usage capture. Do not combine openAIClient with provider constructor connection options such as apiKey or baseURL; configure those on the client.

Handoffs and tool calls inherit scope when they stay in the same execution, but every separately configured model/client must also be instrumented. Remote tools and hosted model calls hidden behind a service are outside the injected client's coverage. Responses WebSocket transport bypasses HTTP responses.create and is not covered by this integration. Realtime has separate trackRealtimeEvent/wrapRealtime helpers and is not covered by the ordinary runner example. Streaming agents must be consumed and their completion awaited inside the scope before the serverless request exits.

Reference: OpenAI Agents SDK repository and documentation.

LangGraph and LangChain

Use createLangChainCallback on the graph or agent invocation. LangGraph inherits LangChain model callbacks, so the handler captures model completions while chain and tool totals do not create additional spend.

import { createAgent } from "langchain";
import { ChatOpenAI } from "@langchain/openai";
import { createLangChainCallback, withMetadata } from "@superpenguin/js";

const agent = createAgent({
  model: new ChatOpenAI({ model: "gpt-4o-mini" }),
  tools: [],
});
const callback = createLangChainCallback({
  metadata: { feature: "support_agent", environment: "production" },
});

const result = await withMetadata(
  { customer_id: "acme" },
  () => agent.invoke(
    { messages: [{ role: "user", content: "How do I reset my password?" }] },
    { callbacks: [callback], metadata: { ticket_id: "123" } },
  ),
);

Invocation metadata is dynamic too. You can supply customer_id there instead of using a scope; the scope is useful when tools call additional wrapped clients outside LangChain. The handler copies scalar metadata and custom_tags, then adds model-run identifiers. Reused callbacks isolate concurrent calls by LangChain run ID.

OpenAI and Anthropic usage shapes are tested. Other LangChain providers may supply compatible standardized usage_metadata, but are not verified by this release. Set explicit provider and model when model identity cannot be inferred. A model that supplies no usage cannot be costed reliably. ChatOpenAI clients routed through OpenRouter, Vercel AI Gateway or LiteLLM require an explicit billing provider override; the callback cannot inspect their base URL or guarantee gateway-reported costs. Use a native client integration when those cost details are required. Framework callbacks report what the framework exposes, rather than hidden provider retry attempts. Native wrappers likewise observe the completed SDK call rather than each internal HTTP retry. Token cache creation without a TTL split uses the existing five-minute cache-write bucket; exact one-hour pricing needs provider-specific usage details.

Do not also wrap a model's underlying client when the callback reports the same invocation. Streaming requires provider usage chunks and complete consumption; callbacks that never receive an end/error event cannot submit a completed usage event. LangGraph checkpoints do not persist the SuperPenguin execution scope; reestablish it when resuming.

References: LangChain agents and LangGraph overview.

Pi

Wrap Pi's model stream function once, then scope metadata around each agent prompt. This targets the current @earendil-works Pi 1.0.4 protocol; legacy @mariozechner packages and CLI subprocesses are not verified.

import { createAgentSession } from "@earendil-works/pi-coding-agent";
import { trackPiStreamFn, withMetadata } from "@superpenguin/js";

const { session } = await createAgentSession();
session.agent.streamFunction = trackPiStreamFn(session.agent.streamFunction, {
  metadata: { feature: "coding_agent" },
});
try {
  await withMetadata({ customer_id: "acme", session_id: "fix-42" }, () =>
    session.prompt("Inspect the failing test, fix the bug, then run the test."),
  );
} finally {
  session.dispose();
}

Each completed model response produces one usage event, including cached input and reported reasoning counts. Pi events and the response object pass through unchanged. Enabled content capture includes the system prompt, visible response, tool arguments, and tool results sent in subsequent model inputs. The tests run a real Pi Agent with a deterministic two-call tool loop; they do not launch the coding CLI. Avoid also instrumenting Pi's provider transport for the same calls. See Pi SDK documentation.

OpenCode

OpenCode requires a plugin inside its server process. A remote SDK client cannot install that hook, and the caller's withMetadata scope cannot cross the process boundary. Use a server-side session resolver for dynamic business metadata.

Create .opencode/plugins/superpenguin.ts with this single plugin export:

import { createOpenCodePlugin } from "@superpenguin/js";
import { lookupSessionMetadata } from "../../session-metadata";

export const SuperPenguinPlugin = createOpenCodePlugin({
  metadata: { feature: "coding_agent" },
  metadataForSession: async (sessionID) => ({
    ...await lookupSessionMetadata(sessionID),
    session_id: sessionID,
  }),
});

lookupSessionMetadata is your application function, for example a database lookup returning { customer_id: "acme" }. Configure SuperPenguin credentials in the server environment and install @superpenguin/js where the plugin can import it. Do not configure the entire SDK package as an OpenCode plugin: its public exports are helpers, not plugin factories.

The adapter targets the stable @opencode-ai/plugin 1.18.34 contract. It reports completed step-finish usage once per model step, reconstructs cache and reasoning totals from OpenCode's normalized fields, and ignores aggregate assistant totals. The plugin supports static defaults and metadata resolved for each model request. It currently captures usage and attribution only; full system prompts, tool contents, and visible answers are not implemented for this adapter. OpenCode's local cost estimate is not treated as a provider-reported charge. Missing cache-write TTL detail uses the existing five-minute bucket.

Compatibility tests use the installed stable plugin types and source-derived events; no OpenCode server is launched. The V2 preview plugin API is not supported. OpenCode discards the global event hook's promise, so the server must remain alive long enough to deliver telemetry. OpenCode awaits the plugin's dispose() hook during orderly instance shutdown; it drains pending telemetry. Forced process exit can still interrupt delivery. See OpenCode plugin documentation.

A bounded analytics workflow

The repository example at sdk/js/examples/analytics-agent.ts runs a real Vercel ToolLoopAgent with four deterministic model responses and local ingest transport:

  1. Inspect the orders schema.
  2. Query September revenue.
  3. Run period comparison and enterprise-segment tools in the same model turn.
  4. Return a concise answer supported by the tool results.

Run it from sdk/js with npx tsx examples/analytics-agent.ts. It needs no provider credentials and makes no live provider calls. Its output is:

Agent answer: September revenue was $1,120, up 12% from August. Enterprise contributed $720.
Provider calls: 4 input tokens: 550 output tokens: 110
Per-call reasoning tokens: [ 4, 6, 8, 10 ]
Locally captured content samples: 4

The workflow tests exercise actual framework instances over fake transports. They assert parallel tool execution, system instructions, tool arguments and results, the final answer, customer metadata, and exactly one usage event per model call. OpenAI Agents and LangChain have additional tool-loop content tests; middleware streaming tests assert original chunks are preserved. These tests verify instrumentation, not a live model's ability to choose the right tools.

Content and reasoning usage

Content capture remains opt-in, controlled by the organization's capture configuration and sampling. Vercel model middleware, LangChain callbacks and Pi stream wrappers normalize their framework message shapes into the existing capture contract. System prompts and the last user turn form prompt_text; tool calls and tool results have separate fields; visible output forms output_text. This is sampled model-call capture, not a complete persisted agent transcript or a trace of every tool lifecycle event. A final tool result that is never sent to another model call is outside the model-input capture path.

Vercel streaming accumulates bounded visible text and completed tool calls until the usage-bearing finish part. Cancellation before that part cannot produce a completed usage/content sample. Raw thinking and encrypted reasoning are excluded from captured outcomes. Native OpenAI Responses can retain provider-supplied visible reasoning summaries using its existing capture path.

When a provider reports a positive reasoning-token count, the adapters preserve it in custom_tags.reasoning_tokens. Reasoning is already included in the output-token total, so the tag is never added as another billing leg. Zero or unavailable reasoning counts do not introduce a tag. OpenAI Responses/Chat Completions, Vercel normalized usage, LangChain standardized usage, Pi, and OpenCode step usage are supported; a provider that does not expose a separate count cannot be reconstructed from the answer text.

Custom harnesses and private agent SDKs

A custom SDK is supported if it gives you access to at least one model-call boundary:

  1. Inject a wrapped OpenAI or Anthropic client.
  2. Inject an instrumented Vercel model.
  3. Install a model-completion callback that reports identity and usage.

For the third option, normalize the SDK's usage into trackUsage. This explicit API inherits the current metadata scope:

import { trackUsage, withMetadata } from "@superpenguin/js";

await withMetadata({ customer_id: "acme", feature: "support_agent" }, async () => {
  await harness.run({
    prompt: "Help with my account",
    onModelComplete: async (call) => {
      await trackUsage({
        provider: call.billingProvider,
        model: call.model,
        usage: { inputTokens: call.inputTokens, outputTokens: call.outputTokens },
        idempotencyKey: call.id,
      });
    },
  });
});

harness.run and onModelComplete are illustrative hooks, not a universal SDK interface. The custom SDK must actually expose equivalent data. Preserve cached-token and other usage details when available, and use an idempotency key unique to the invocation. The normalized inputTokens includes cache reads; Anthropic cache creation is excluded from inputTokens and supplied separately as cacheWrite5mTokens/cacheWrite1hTokens. OpenAI input counts retain cache writes. SuperPenguin cannot introspect a closed SDK that exposes no client, model hook, usable usage result or supported trace. If only an aggregate run total is available, per-model costs cannot be reconstructed accurately for mixed-model runs.

Python already supports with sp.metadata({...}) for wrapped provider calls. These new adapters and examples target TypeScript; this release does not add a Python LangChain callback or a Vercel Python integration.