Skip to content

[ai-gateway-provider] Concurrent requests can receive each other's responses #620

Description

@MaratFM

Summary

When an AiGatewayChatLanguageModel is created once and reused for concurrent requests, a request can receive a response generated for another request. This can expose one user's model output to another user in applications that share a model instance.

Internally, processModelRequest() replaces each wrapped model's config.fetch with a request-local collector. Concurrent calls mutate the same shared config.fetch, causing one collector to capture multiple requests while another captures none. The resulting gateway responses can then be missing or returned to the wrong caller.

Environment

  • ai-gateway-provider@3.2.0
  • ai@6.0.208
  • Node.js 24.12.0

Reproduction

mkdir ai-gateway-concurrency-repro
cd ai-gateway-concurrency-repro
npm init -y
npm install ai@6.0.208 ai-gateway-provider@3.2.0

Save as repro.mjs:

import { generateText } from 'ai';
import { createAiGateway } from 'ai-gateway-provider';
import { createDeepSeek } from 'ai-gateway-provider/providers/deepseek';

const gateway = createAiGateway({
  accountId: 'test',
  gateway: 'test',
  apiKey: 'test',
});

const deepseek = createDeepSeek();
const model = gateway(deepseek('deepseek-chat'));
const gatewayBodies = [];

globalThis.fetch = async (_url, init) => {
  const body = JSON.parse(init.body);
  const prompts = body.map(request => request.query?.messages?.at(-1)?.content);
  gatewayBodies.push(prompts);

  const prompt = prompts[0];
  return new Response(
    JSON.stringify({
      id: 'id',
      object: 'chat.completion',
      created: 0,
      model: 'deepseek-chat',
      choices: [
        {
          index: 0,
          message: { role: 'assistant', content: `ANSWER:${prompt}` },
          finish_reason: 'stop',
        },
      ],
      usage: {
        prompt_tokens: 1,
        completion_tokens: 1,
        total_tokens: 2,
      },
    }),
    {
      headers: {
        'content-type': 'application/json',
        'cf-aig-step': '0',
      },
    },
  );
};

const ask = prompt => generateText({ model, prompt, maxRetries: 0 }).then(result => result.text);

const [a, b] = await Promise.all([ask('A'), ask('B')]);

console.log({ a, b, gatewayBodies });

Run:

node repro.mjs

Output:

{
  a: 'ANSWER:undefined',
  b: 'ANSWER:A',
  gatewayBodies: [ [], [ 'A', 'B' ] ]
}

The reproduction stubs only the final gateway fetch; no Cloudflare or provider credentials are required.

Expected behavior

Concurrent calls using the same model remain isolated:

{ a: 'ANSWER:A', b: 'ANSWER:B', gatewayBodies: [ [ 'A' ], [ 'B' ] ] }

Actual behavior

One gateway request has an empty body and the other contains both provider requests. The responses are consequently missing or assigned to the wrong caller.

Workaround

Create the provider model and gateway wrapper separately for each call instead of sharing the resulting model instance.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions