Two dark teal glacial streams converge into one channel between textured white and pale blue ice banks, seen from directly above
All posts

How to make LLM tool calls idempotent so a retry doesn't create a duplicate ticket

Deduplicate the same tool operation; changed arguments create a new one. Require an explicit conflict policy to allow only one write per request.

Deduplicate repeated LLM calls to the same operation by deriving a backend key from the authenticated tenant, a stable application request_id, the function name, and canonical arguments. Changed arguments create a new operation; allowing only one write per request requires an explicit conflict policy. Claim the key atomically with a UNIQUE constraint, commit the local operation and result together, and return that result on repeats. External destinations must also deduplicate or support reliable reconciliation; a local key cannot close that atomicity gap.

This guide implements that boundary in an AIVAX Protocol Function callback using Bun and SQLite. It does not promise exactly-once execution across services.

Why put the transaction boundary in your backend?

A ticket can exist even when its confirmation never reaches the caller. An error response therefore leaves the effect uncertain. A model instruction cannot provide the durable record and lock needed to retry safely.

Tian Pan frames tool-call idempotency as a distributed-systems problem; the DEV duplicate-delivery walkthrough uses a ledger lookup; Retell warns against double-booking after conversational interruptions. These sources motivate the design without measuring AIVAX duplicate rates.

Which AIVAX fields identify the operation?

A Protocol Function makes one HTTP POST per invocation, with no internal retry in that callback path. It sends function.name, model-generated arguments in function.content, and a separately constructed context containing externalUserId, metadata, callSource, conversationToken, and moment. It does not supply a unique callback ID.

Your application backend should create metadata.request_id for a logical operation and preserve it across retries, whether supplied through request or session metadata. Change it when the user intentionally starts another operation. A session-wide value reused indefinitely would suppress later legitimate identical actions. A client-chosen identifier without authentication proves neither identity nor permission.

Resolve externalUserId against your application's authorized user records, then obtain the tenant from that record. Do not select the tenant from model arguments or accept metadata.tenant_id as proof of access. The downloadable demo uses a small configured user map; production needs an authenticated session or request binding, resource authorization, and a check that the operation belongs to that caller. Authenticate and authorize before returning a cached result, too.

Neither conversationToken nor moment belongs in this key. The former correlates a conversation and can be shared by several calls; it is not a callback ID. The latter changes when another callback is prepared. The same separation between identity metadata and enforced permission underlies our MCP trust-boundary guide.

The account nonce answers a different question. When an account has a hook key, AIVAX sends X-Request-Nonce: a variable BCrypt hash of that key, generated with work factor 9. A production callback verifies the received hash against the hook key with BCrypt verification, rather than comparing it with one fixed hash. This verifies the account hook key, not individual user permissions; it neither signs the body nor prevents replay. Keep the endpoint on a protected, authenticated transport and enforce your own operation controls.

Derive a stable key without confusing different intents

The excerpts below require the complete tested server.js, which supplies the HTTP handler, local access controls, validation, and logging.

The callback accepts a title, an optional description, and optional object details. After validation, it recursively sorts object keys while preserving array order:

function canonicalJson(value, depth = 0) {
  if (depth > 64 || (typeof value === "number" && !Number.isFinite(value))) {
    throw new TypeError("JSON exceeds the allowed limits.");
  }
  if (Array.isArray(value)) {
    return `[${value.map(item => canonicalJson(item, depth + 1)).join(",")}]`;
  }
  if (value !== null && typeof value === "object") {
    return `{${Object.keys(value).sort().map(
      key => `${JSON.stringify(key)}:${canonicalJson(value[key], depth + 1)}`,
    ).join(",")}}`;
  }
  return JSON.stringify(value);
}

The request handler hashes an unambiguous JSON tuple, using createHash from node:crypto:

const key = createHash("sha256").update(JSON.stringify([
  user.tenantId, requestId, body.function.name, argumentsJson,
])).digest("hex");

Here, argumentsJson is the canonicalized argument object. This is deterministic for the accepted JSON values, but it is not a complete RFC 8785 implementation. It does not make paraphrases equivalent: changing a title's wording or an array's order changes the key.

Different arguments under the same request_id permit a corrected request, but also a semantically duplicate ticket with different wording. If only one ticket is allowed per approved request, make tenant plus request ID plus function unique, store the argument fingerprint separately, and reject changed arguments. That stricter conflict policy requires changing the example; the tested implementation remains argument-sensitive.

Claim the key and create the local ticket in one transaction

A process-local Set loses its memory on restart and cannot coordinate two server processes. SQLite supplies the durable constraint here. The operations.idempotency_key primary key is unique, and the ticket table also has operation_key TEXT NOT NULL UNIQUE REFERENCES operations(idempotency_key). The store uses WAL mode, foreign keys, full synchronous writes, and a bounded busy timeout.

The transaction first claims the key. A conflict returns the stored result rather than repeating the insert:

const createTicket = db.transaction((key, user, argumentsJson) => {
  const claimed = db.query(`
    INSERT INTO operations (idempotency_key, user_id, state)
    VALUES (?, ?, 'pending') ON CONFLICT(idempotency_key) DO NOTHING
  `).run(key, user.userId);

  if (!claimed.changes) {
    const original = db.query("SELECT * FROM operations WHERE idempotency_key = ?").get(key);
    if (original.user_id !== user.userId) {
      return reject("request_id_conflict", "request_id already belongs to another user in this tenant.", 403);
    }
    if (original.state === "pending") {
      return Response.json({ status: "in_progress", retryable: true, message: "Operation pending; do not create another one. Check again with the same data." });
    }
    return Response.json({ ...JSON.parse(original.result_json), deduplicated: true });
  }

  const result = { ticket_id: randomUUID(), status: "created" };
  db.query(`
    INSERT INTO tickets (ticket_id, operation_key, tenant_id, user_id, arguments_json)
    VALUES (?, ?, ?, ?, ?)
  `).run(result.ticket_id, key, user.tenantId, user.userId, argumentsJson);
  db.query("UPDATE operations SET state = 'completed', result_json = ? WHERE idempotency_key = ?")
    .run(JSON.stringify(result), key);
  return Response.json({ ...result, deduplicated: false });
});

The handler calls createTicket.immediate(key, user, argumentsJson), so SQLite acquires the write transaction before the claim. The operation row, ticket, and saved result commit together. An exception inside that transaction rolls them back. A later response or logging failure cannot undo a committed ticket; retry with the same key to recover its result. SQLite coordinates both processes sharing this local database.

The repeat preserves the original ticket_id and status; only the delivery annotation deduplicated changes. A retained pending record returns in_progress without taking over the operation. The tests inject that state explicitly: normal successful local execution commits the completed result in the same transaction. Reconcile unresolved operations before considering another execution.

What should the model see when a callback rejects a call?

For a non-2xx response or an exception, AIVAX passes exactly “The tool call resulted in an error.” to the model. Your non-2xx response body is not passed through; diagnostic details go to the log. The model may call again, but that is not guaranteed and is not an internal HTTP retry.

Use HTTP 200 with a structured rejection for correctable input errors the model needs to understand. Keep authentication and authorization failures as 401/403, and internal failures as 500. Returning 200 does not mean the business action succeeded: the response body must state its status clearly.

Callback responses and the information delivered to the model
SituationCallback responseWhat the model seesBackend action
Correctable validation failureHTTP 200, {status: "rejected", error: {code, message}}The rejection body, including the correction neededDo not create a ticket
Authentication or authorization failureHTTP 401 or 403The tool call resulted in an error.Reject before any write or cached-result disclosure
Internal failureHTTP 500The tool call resulted in an error.Roll back uncommitted writes; after commit, recover the outcome with the same key
Retained pending operationHTTP 200, status: "in_progress"The pending status and instruction not to create another operationReturn pending; reconcile rather than re-execute
SQLite writer unavailableHTTP 200, status: "busy", retryable: trueA retryable busy resultRetry the same operation later, with a bounded policy
Completed duplicateHTTP 200, stored result plus deduplicated: trueThe original ticket ID and outcomeNo new ticket

For example, the demo rejects a missing externalUserId with a structured HTTP 200 rejection and no write, while an unknown user receives 403. In production, the application must supply authenticated identity; the assistant must not invent one to satisfy the error. retryable and these status names are application conventions. Configure bounded retries and escalation in your application; these fields do not schedule retries. Our guide to state in stateless MCP covers the related responsibility for durable application state and retries.

Return only the status and fields needed for the next step, without internal traces. The callback response has a 20 MB read limit.

Run two identical calls and inspect the stored result

Download server.js and config.json.example into an empty local directory. With Bun installed, copy the configuration and start the service:

cp config.json.example config.json
bun server.js --config config.json

The configuration uses tickets.sqlite, port 3210, and a synthetic user-a mapped to tenant-a and user alice. It binds only to 127.0.0.1. The fixed nonce only simulates authentication locally. Do not expose this demo as a production callback. Real deployment requires BCrypt verification, protected transport, and application-bound identity and request authorization.

In another terminal, send the same payload twice:

for attempt in 1 2; do
  curl --fail-with-body --silent --show-error \
    http://127.0.0.1:3210/callback \
    -H 'Content-Type: application/json' \
    -H 'X-Request-Nonce: local-demo-nonce-0123456789' \
    --data '{
      "type": "tool_call",
      "function": {
        "name": "create_ticket",
        "content": {"title": "Cannot sign in"}
      },
      "context": {
        "externalUserId": "user-a",
        "metadata": {"request_id": "request-example"}
      }
    }'
  printf '\n'
done

On a fresh database, the first response has status: "created" and deduplicated: false. The second returns the same generated ticket_id with deduplicated: true. Running the pair again against the same database returns the stored ticket both times. Inspect debug.log for the operation key and response status; it does not log the full payload.

Local checks used Bun 1.3.14, curl, and SQLite, without inference or production benchmarking. Ten concurrent calls across two processes sharing one database produced one creation and nine deduplicated responses. Checks also confirmed argument-sensitive keys, object-key-order independence, tenant isolation, nonce rejection, identity and argument validation, authorization, busy under a write lock, and no re-execution of injected pending records. An injected insert failure rolled back both writes; a server restart reused the saved result. Not tested: abrupt process death during commit, real .NET callback timeout behavior, real BCrypt authentication, end-to-end inference, or a distributed store.

How do you keep an external ticket system from duplicating the write?

Do not replace the SQLite insert with a CRM request and assume the transaction still protects it. If the remote system creates the ticket and the callback dies before saving the result, the next attempt can create a second ticket. Recording the key first merely changes the failure into an unresolved pending operation; it does not reveal whether the remote write happened.

When the destination supports idempotency keys, forward the same stable key, persist the returned ticket ID, and reconcile uncertain outcomes using the destination's contract. Match its key scope and retention window. If it supports lookup by a unique operation identifier, use that to resolve uncertainty, and verify its uniqueness and consistency guarantees before allowing another write.

For asynchronous execution, persist the operation and an outbox record in the same transaction. A worker delivers the write with bounded retries and saves or reconciles the result. An outbox alone does not prevent duplicate external effects. The destination must deduplicate or provide a reliable unique-ID lookup and reconciliation path. Without either, stop and surface uncertainty rather than promising exactly-once behavior.

Retention also defines the deduplication window. Deleting an operation key makes its next delivery look new. Keep keys and results for the retry and reconciliation horizon your business needs, and restrict access to stored results as carefully as access to the original ticket.

What else should you check before production?

Can I use conversationToken instead of request_id?

No: it may be shared by multiple calls or null. Keep a backend-issued operation identifier stable across retries.

Does in_progress mean it is safe to try a new key?

No: a new key bypasses the unresolved record. Preserve the key, check authoritative state, and escalate if necessary.

Review the Protocol Functions contract and test duplicate delivery and ambiguous outcomes against the actual destination before enabling production writes.