Skip to content

feat(aiguard): evaluating anthropic calls with AI guard automatically - #9563

Open
IlyasShabi wants to merge 5 commits into
masterfrom
ishabi/anthropic-aiguard
Open

feat(aiguard): evaluating anthropic calls with AI guard automatically#9563
IlyasShabi wants to merge 5 commits into
masterfrom
ishabi/anthropic-aiguard

Conversation

@IlyasShabi

@IlyasShabi IlyasShabi commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

What does this PR do?

This PR adds automatic AI Guard integration for the Anthropic SDK via auto-instrumentation. Inside the wrapped parse() and asResponse() calls, we evaluate the input (before the model runs) and the input/output (after the model responds) and hand them off to AI Guard for evaluation. If AI Guard denies, the call is aborted before the response reaches user.

Adds a new integration file packages/dd-trace/src/aiguard/integrations/anthropic.js that subscribes to the Anthropic lifecycle channels and normalizes anthropicmessages into AI Guard's style.

Streaming is out of scope.

Motivation

AI Guard already integrates with the OpenAI SDK and Vercel AI SDK. Anthropic is also a major LLM provider in the tracer supported set without AI Guard coverage.

Additional Notes

JIRA: https://datadoghq.atlassian.net/browse/APPSEC-62251

…#9219)

* feat(aiguard): evaluating anthropic calls with AI guard automatically
@dd-octo-sts

dd-octo-sts Bot commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Overall package size

Self size: 7.61 MB
Deduped: 8.27 MB
No deduping: 8.27 MB

Dependency sizes | name | version | self size | total size | |------|---------|-----------|------------| | import-in-the-middle | 3.3.2 | 124.41 kB | 440.65 kB | | opentracing | 0.14.7 | 194.81 kB | 194.81 kB | | dc-polyfill | 0.1.11 | 25.74 kB | 25.74 kB |

🤖 This report was automatically generated by heaviest-objects-in-the-universe

@datadog-datadog-prod-us1

datadog-datadog-prod-us1 Bot commented Jul 28, 2026

Copy link
Copy Markdown

Tests

🎉 All green!

🧪 All tests passed
❄️ No new flaky tests detected

🔄 Datadog retried 1 test - 1 passed on retry View in Datadog

🎯 Code Coverage (details)
Patch Coverage: 100.00%
Overall Coverage: 98.51% (+0.01%)

This comment will be updated automatically if new data arrives.
🔗 Commit SHA: cabca2b | Docs | Datadog PR Page | Give us feedback!

@codecov

codecov Bot commented Jul 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.51%. Comparing base (fdd7d8b) to head (cabca2b).
⚠️ Report is 15 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #9563      +/-   ##
==========================================
+ Coverage   98.50%   98.51%   +0.01%     
==========================================
  Files         952      956       +4     
  Lines      131108   132480    +1372     
  Branches    11128    11425     +297     
==========================================
+ Hits       129145   130515    +1370     
- Misses       1963     1965       +2     
Flag Coverage Δ
aiguard 57.57% <100.00%> (+0.17%) ⬆️
aiguard-integration 55.71% <42.12%> (-0.51%) ⬇️
apm-bucket-0 57.31% <ø> (-0.37%) ⬇️
apm-bucket-1 63.49% <ø> (-0.40%) ⬇️
apm-bucket-2 62.24% <ø> (-0.41%) ⬇️
apm-bucket-3 59.74% <ø> (-0.38%) ⬇️
apm-capabilities-tracing 62.37% <42.88%> (-0.32%) ⬇️
apm-integrations-aerospike 56.33% <ø> (-0.36%) ⬇️
apm-integrations-confluentinc-kafka-javascript 61.12% <ø> (-0.41%) ⬇️
apm-integrations-couchbase 56.74% <ø> (-0.36%) ⬇️
apm-integrations-http 62.30% <ø> (-0.38%) ⬇️
apm-integrations-kafkajs 61.72% <ø> (-0.41%) ⬇️
apm-integrations-next 58.80% <ø> (-0.37%) ⬇️
apm-integrations-prisma 58.36% <ø> (-0.31%) ⬇️
appsec 72.27% <ø> (-0.41%) ⬇️
appsec-express_fastify_graphql 69.69% <ø> (-0.55%) ⬇️
appsec-integration 50.66% <23.88%> (-0.24%) ⬇️
appsec-kafka_ldapjs_lodash 63.49% <ø> (-0.38%) ⬇️
appsec-mongodb-core_mongoose_mysql 67.20% <ø> (-0.37%) ⬇️
appsec-next 57.24% <ø> (-0.30%) ⬇️
appsec-node-serialize_passport_postgres 66.84% <ø> (-0.40%) ⬇️
appsec-sourcing_stripe_template 65.23% <ø> (-0.37%) ⬇️
debugger 64.38% <ø> (-0.39%) ⬇️
instrumentations-bucket-0 51.49% <100.00%> (-0.33%) ⬇️
instrumentations-bucket-1 59.76% <ø> (-0.44%) ⬇️
instrumentations-bucket-10 61.33% <ø> (-0.66%) ⬇️
instrumentations-bucket-11 56.43% <ø> (+4.70%) ⬆️
instrumentations-bucket-12 51.75% <ø> (-0.55%) ⬇️
instrumentations-bucket-13 51.78% <ø> (-0.07%) ⬇️
instrumentations-bucket-2 53.06% <ø> (-0.65%) ⬇️
instrumentations-bucket-3 53.34% <ø> (-5.88%) ⬇️
instrumentations-bucket-4 58.80% <ø> (+6.42%) ⬆️
instrumentations-bucket-5 49.88% <ø> (-7.71%) ⬇️
instrumentations-bucket-6 60.48% <ø> (-0.27%) ⬇️
instrumentations-bucket-7 57.98% <ø> (-0.45%) ⬇️
instrumentations-bucket-8 59.05% <ø> (-0.50%) ⬇️
instrumentations-bucket-9 51.75% <ø> (-9.77%) ⬇️
instrumentations-instrumentation-couchbase 50.77% <ø> (-0.31%) ⬇️
instrumentations-instrumentation-zlib 51.42% <ø> (?)
instrumentations-integration-esbuild 33.99% <23.88%> (-0.03%) ⬇️
llmobs-ai_anthropic_bedrock 62.54% <71.11%> (-0.33%) ⬇️
llmobs-bucket-1 61.29% <60.55%> (-0.32%) ⬇️
llmobs-openai 61.85% <ø> (-0.35%) ⬇️
llmobs-openai-agents_vertex-ai 59.80% <ø> (-0.32%) ⬇️
llmobs-sdk 66.33% <ø> (+0.35%) ⬆️
master-coverage 98.51% <100.00%> (?)
openfeature 55.76% <ø> (-0.32%) ⬇️
openfeature-unit 53.12% <ø> (-0.32%) ⬇️
platform-core_esbuild_instrumentations-misc 40.67% <23.88%> (-0.12%) ⬇️
platform-integration 60.69% <ø> (-0.36%) ⬇️
platform-shimmer_unit-guardrails_webpack 38.93% <23.88%> (-0.13%) ⬇️
plugins-bucket-0 56.78% <ø> (-0.31%) ⬇️
plugins-bucket-1 54.02% <ø> (-0.35%) ⬇️
plugins-bucket-11 61.87% <ø> (-0.37%) ⬇️
plugins-bucket-17 ?
plugins-bucket-18 61.57% <ø> (-0.86%) ⬇️
plugins-bucket-19 59.66% <ø> (-2.00%) ⬇️
plugins-bucket-20 61.64% <ø> (-2.59%) ⬇️
plugins-bucket-4 58.23% <ø> (-0.36%) ⬇️
plugins-bullmq_cassandra_cookie 61.32% <ø> (-0.39%) ⬇️
plugins-cookie-parser_crypto_dd-trace-api 56.39% <ø> (-0.36%) ⬇️
plugins-fetch_fs_generic-pool 58.44% <ø> (-0.36%) ⬇️
plugins-google-cloud-pubsub_grpc_handlebars 64.22% <ø> (-0.41%) ⬇️
plugins-hapi_hono_ioredis 59.84% <ø> (-0.37%) ⬇️
plugins-jest_knex_langgraph 55.30% <ø> (-0.34%) ⬇️
plugins-ldapjs_light-my-request_limitd-client 58.15% <ø> (-0.36%) ⬇️
plugins-lodash_mariadb_memcached 57.73% <ø> (-0.37%) ⬇️
plugins-moleculer_mongodb_mongodb-core 61.49% <ø> (-0.39%) ⬇️
plugins-mongoose_multer_mysql 58.73% <ø> (-0.36%) ⬇️
plugins-mysql2_nats_node-serialize 60.30% <ø> (-0.39%) ⬇️
plugins-opensearch_passport-http_pino 59.19% <ø> (-0.37%) ⬇️
plugins-postgres_process_pug 58.33% <ø> (?)
plugins-process_pug_redis ?
plugins-redis_router_sequelize 61.69% <ø> (?)
plugins-test-and-upstream-rhea_undici_url 61.27% <ø> (?)
plugins-undici_url_valkey ?
plugins-valkey_vm_winston 57.67% <ø> (?)
plugins-vm_winston_ws ?
plugins-ws 59.21% <ø> (?)
profiling 61.74% <ø> (-0.37%) ⬇️
serverless-aws-sdk-aws-sdk 54.85% <ø> (-0.28%) ⬇️
serverless-aws-sdk-base-inject-field 50.71% <ø> (-0.30%) ⬇️
serverless-aws-sdk-bedrockruntime 54.52% <ø> (-0.30%) ⬇️
serverless-aws-sdk-client 56.16% <ø> (-0.33%) ⬇️
serverless-aws-sdk-dynamodb 55.40% <ø> (-0.32%) ⬇️
serverless-aws-sdk-eventbridge 49.27% <ø> (-0.24%) ⬇️
serverless-aws-sdk-kinesis 59.00% <ø> (-0.34%) ⬇️
serverless-aws-sdk-lambda 57.11% <ø> (-0.32%) ⬇️
serverless-aws-sdk-s3 55.50% <ø> (-0.31%) ⬇️
serverless-aws-sdk-serverless-peer-service 59.41% <ø> (-0.35%) ⬇️
serverless-aws-sdk-sns 59.84% <ø> (-0.35%) ⬇️
serverless-aws-sdk-sqs 60.27% <ø> (-0.36%) ⬇️
serverless-aws-sdk-stepfunctions 55.33% <ø> (-0.31%) ⬇️
serverless-aws-sdk-util 51.27% <ø> (-0.31%) ⬇️
serverless-bucket-0 53.90% <ø> (-0.35%) ⬇️
serverless-bucket-1 58.91% <ø> (-0.34%) ⬇️
test-optimization-cucumber 71.58% <ø> (-0.43%) ⬇️
test-optimization-cypress 65.44% <ø> (-0.30%) ⬇️
test-optimization-jest 72.87% <ø> (-0.43%) ⬇️
test-optimization-mocha 72.78% <ø> (-0.37%) ⬇️
test-optimization-playwright-playwright-atr 60.21% <ø> (-0.35%) ⬇️
test-optimization-playwright-playwright-efd 60.40% <ø> (-0.34%) ⬇️
test-optimization-playwright-playwright-final-status 60.52% <ø> (-0.35%) ⬇️
test-optimization-playwright-playwright-impacted-tests 60.10% <ø> (-0.19%) ⬇️
test-optimization-playwright-playwright-reporting 61.31% <ø> (-0.46%) ⬇️
test-optimization-playwright-playwright-test-management 60.90% <ø> (-0.46%) ⬇️
test-optimization-playwright-playwright-test-span 60.31% <ø> (-0.41%) ⬇️
test-optimization-selenium 59.81% <ø> (-0.47%) ⬇️
test-optimization-testopt 58.15% <ø> (-0.17%) ⬇️
test-optimization-vitest 70.34% <ø> (-0.34%) ⬇️
test-optimization-vitest-browser 59.05% <ø> (-0.23%) ⬇️
test-optimization-webdriverio 64.27% <ø> (-0.29%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@IlyasShabi
IlyasShabi marked this pull request as ready for review July 28, 2026 13:19
@IlyasShabi
IlyasShabi requested review from a team as code owners July 28, 2026 13:19
@IlyasShabi
IlyasShabi requested review from tlhunter and removed request for a team July 28, 2026 13:19

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 199b4e4c53

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/dd-trace/src/aiguard/messages/anthropic.js Outdated
Comment thread packages/datadog-instrumentations/src/anthropic.js Outdated
Comment thread packages/datadog-instrumentations/src/anthropic.js Outdated
@pr-commenter

pr-commenter Bot commented Jul 28, 2026

Copy link
Copy Markdown

Benchmarks

Benchmark execution time: 2026-07-31 08:51:16

Comparing candidate commit cabca2b in PR branch ishabi/anthropic-aiguard with baseline commit fdd7d8b in branch master.

📊 Benchmarking dashboard

Found 0 performance improvements and 0 performance regressions! Performance is the same for 2325 metrics, 33 unstable metrics.

Explanation

This is an A/B test comparing a candidate commit's performance against that of a baseline commit. Performance changes are noted in the tables below as:

  • 🟩 = significantly better candidate vs. baseline
  • 🟥 = significantly worse candidate vs. baseline

We compute a confidence interval (CI) over the relative difference of means between metrics from the candidate and baseline commits, considering the baseline as the reference.

If the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD), the change is considered significant.

Feel free to reach out to #apm-benchmarking-platform on Slack if you have any questions.

More details about the CI and significant changes

You can imagine this CI as a range of values that is likely to contain the true difference of means between the candidate and baseline commits.

CIs of the difference of means are often centered around 0%, because often changes are not that big:

---------------------------------(------|---^--------)-------------------------------->
                              -0.6%    0%  0.3%     +1.2%
                                 |          |        |
         lower bound of the CI --'          |        |
sample mean (center of the CI) -------------'        |
         upper bound of the CI ----------------------'

As described above, a change is considered significant if the CI is entirely outside the configured SIGNIFICANT_IMPACT_THRESHOLD (or the deprecated UNCONFIDENCE_THRESHOLD).

For instance, for an execution time metric, this confidence interval indicates a significantly worse performance:

----------------------------------------|---------|---(---------^---------)---------->
                                       0%        1%  1.3%      2.2%      3.1%
                                                  |   |         |         |
       significant impact threshold --------------'   |         |         |
                      lower bound of CI --------------'         |         |
       sample mean (center of the CI) --------------------------'         |
                      upper bound of CI ----------------------------------'

Unstable benchmarks

These benchmarks have a confidence interval too wide to call a change; treat them as noise rather than signal.

scenario:appsec-appsec-enabled-24

  • unstable execution_time [-208.621ms; +203.837ms] or [-7.850%; +7.670%]

scenario:appsec-appsec-enabled-26

  • unstable execution_time [-241.855ms; +239.730ms] or [-9.443%; +9.360%]

scenario:appsec-appsec-enabled-with-attacks-24

  • unstable execution_time [-165.316ms; +158.316ms] or [-5.352%; +5.125%]

scenario:appsec-appsec-enabled-with-attacks-26

  • unstable execution_time [-191.425ms; +183.787ms] or [-6.582%; +6.319%]

scenario:appsec-control-20

  • unstable execution_time [-130595.822µs; +129119.455µs] or [-7.966%; +7.876%]

scenario:appsec-control-24

  • unstable execution_time [-114062.411µs; +114913.911µs] or [-9.189%; +9.257%]

scenario:appsec-control-26

  • unstable execution_time [-124498.104µs; +125257.937µs] or [-10.038%; +10.100%]

scenario:appsec-iast-no-vulnerability-control-20

  • unstable execution_time [-14.111ms; +27.610ms] or [-5.358%; +10.484%]

scenario:appsec-iast-no-vulnerability-iast-enabled-always-active-20

  • unstable execution_time [-19.560ms; +13.036ms] or [-7.608%; +5.070%]

scenario:appsec-iast-with-vulnerability-control-20

  • unstable execution_time [-34341.842µs; +35890.733µs] or [-6.237%; +6.518%]

scenario:child_process-shell-string-24

  • unstable execution_time [-11.663ms; +22.765ms] or [-3.637%; +7.099%]

scenario:debugger-line-probe-with-snapshot-default-24

  • unstable cpu_user_time [-2049.083ms; +3220.163ms] or [-24.737%; +38.875%]
  • unstable execution_time [-2152.511ms; +3376.740ms] or [-23.983%; +37.623%]
  • unstable instructions [-17.1G instructions; +27.3G instructions] or [-25.277%; +40.383%]
  • unstable max_rss_usage [-8.065MB; +13.370MB] or [-5.142%; +8.524%]
  • unstable throughput [-883.308op/s; +568.898op/s] or [-24.002%; +15.458%]

scenario:debugger-line-probe-with-snapshot-default-26

  • unstable cpu_user_time [-2637.577ms; +4202.940ms] or [-27.599%; +43.979%]
  • unstable execution_time [-2653.680ms; +4214.256ms] or [-25.764%; +40.915%]
  • unstable instructions [-23.6G instructions; +37.3G instructions] or [-29.631%; +46.848%]
  • unstable max_rss_usage [-9.486MB; +13.670MB] or [-5.967%; +8.599%]
  • unstable throughput [-819.879op/s; +518.184op/s] or [-25.464%; +16.094%]

scenario:debugger-line-probe-without-snapshot-24

  • unstable cpu_user_time [-1772.838ms; +571.202ms] or [-21.362%; +6.883%]
  • unstable execution_time [-1777.737ms; +560.594ms] or [-19.754%; +6.229%]
  • unstable instructions [-15.1G instructions; +4.8G instructions] or [-22.260%; +7.116%]
  • unstable throughput [-148.404op/s; +476.834op/s] or [-4.054%; +13.027%]

scenario:dogstatsd-with-tags-20

  • unstable cpu_user_time [-415.918ms; +267.794ms] or [-8.679%; +5.588%]
  • unstable execution_time [-413.949ms; +272.389ms] or [-8.505%; +5.596%]
  • unstable throughput [-98029.439op/s; +146138.712op/s] or [-5.680%; +8.467%]

scenario:fs-subscribed-24

  • unstable execution_time [-15.943ms; +28.352ms] or [-3.965%; +7.052%]

scenario:plugin-graphql-long-with-depth-off-26

  • unstable max_rss_usage [-41.939MB; +22.140MB] or [-20.113%; +10.618%]

scenario:plugin-graphql-long-with-depth-on-max-20

  • unstable cpu_user_time [-591.137ms; +598.580ms] or [-5.114%; +5.178%]
  • unstable execution_time [-602.871ms; +610.905ms] or [-5.113%; +5.181%]
  • unstable throughput [-3.527op/s; +3.511op/s] or [-5.169%; +5.145%]

@tlhunter
tlhunter marked this pull request as draft July 28, 2026 16:04
@tlhunter

Copy link
Copy Markdown
Member

It looks like this is a WIP so I converted it into a draft.

@IlyasShabi

Copy link
Copy Markdown
Contributor Author

@tlhunter I enabled it for codex reviews requests

@IlyasShabi

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d743fc770d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/datadog-instrumentations/src/anthropic.js
Comment thread packages/datadog-instrumentations/src/anthropic.js Outdated
Comment thread packages/dd-trace/src/aiguard/messages/anthropic.js Outdated
@IlyasShabi
IlyasShabi force-pushed the ishabi/anthropic-aiguard branch from d743fc7 to 5bf825e Compare July 29, 2026 13:30
@IlyasShabi

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Swish!

Reviewed commit: 5bf825ecdc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@tlhunter

Copy link
Copy Markdown
Member

@IlyasShabi TIL Codex doesn't review drafts. Maybe we need a better flow then; ideally Codex can perform a review without having GitHub ping folks on slack for a review.

@IlyasShabi

Copy link
Copy Markdown
Contributor Author

Well we can use #9563 (comment) to trigger a review even on draft :D

@IlyasShabi
IlyasShabi force-pushed the ishabi/anthropic-aiguard branch from 46d3c0e to 459ca4e Compare July 30, 2026 15:40
@IlyasShabi

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 459ca4ee6b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread packages/datadog-instrumentations/src/anthropic.js Outdated
Comment thread packages/dd-trace/src/aiguard/messages/anthropic.js Outdated
@IlyasShabi

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: cabca2bc9b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

try {
return finishResult(ctx, JSON.parse(body), getVerdict, body)
} catch {
finish(ctx)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid finalizing nested text reads as success

When asResponse() returns a node-fetch Response, response.json() delegates to this.text(). Because both readers are wrapped, malformed JSON enters the inner text wrapper, this catch calls finish(ctx) as a success, and only afterward the outer json() rejects; its catch skips error publication because ctx.finished is already true. This records anthropic.request as successful even though the reader failed. Avoid finalizing from the nested text call, and cover invalid JSON through the real node-fetch response path.

AGENTS.md reference: AGENTS.md:L128-L129

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not a bug on the supported node-fetch. node-fetch 2.7 consumes the body directly; it never calls this.text().

@IlyasShabi
IlyasShabi marked this pull request as ready for review July 31, 2026 12:17
env:
PLUGINS: anthropic|anthropic-lifecycle
steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0
- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

@tlhunter tlhunter left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Small nit, otherwise LGTM

@BridgeAR BridgeAR left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just blocking since it got an approval and I think the comments are still important

const input = { messages: options.messages }
if (options.system !== undefined) input.system = options.system

const snapshot = [...args]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should move this into the try to make sure copying works, since it is not guaranteed to be an iterable.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

args is guaranteed to be an array because it comes directly from wrapCreate, also make sense to keep it outside and reserve the try/catch for structuredClone failure only.

*/
function snapshotLifecycleArgs (args) {
const options = args[0]
if (!options || typeof options !== 'object') return args

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This keeps the actual arguments and I think we should skip inspection if something is wrong.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will not cause an evaluation, when we have no valid messages array we're going to skip evaluations

Comment on lines +215 to +229
.then(response => {
// Raw output evaluation supports the common json() and text() readers only.
if (!stream &&
(anthropicTracingChannel.start.hasSubscribers ||
afterVerdict ||
messagesAfterChannel.hasSubscribers) &&
wrappedResponse !== response) {
wrappedResponse = response
wrapResponseReader(response, 'json', ctx, getAfterVerdict)
wrapResponseReader(response, 'text', ctx, getAfterVerdict)
}

if (afterVerdict) return afterVerdict.then(() => response)
return response
})

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
.then(response => {
// Raw output evaluation supports the common json() and text() readers only.
if (!stream &&
(anthropicTracingChannel.start.hasSubscribers ||
afterVerdict ||
messagesAfterChannel.hasSubscribers) &&
wrappedResponse !== response) {
wrappedResponse = response
wrapResponseReader(response, 'json', ctx, getAfterVerdict)
wrapResponseReader(response, 'text', ctx, getAfterVerdict)
}
if (afterVerdict) return afterVerdict.then(() => response)
return response
})

I believe this is the issue about asResponse being complained about by the AI findings.

What about removing this for now so that we can land partial support right away and land support for this afterwards as follow-up? :)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added asResponse support in the initial PR and addressed some AI reviews such as #9492 (comment) as you noted.

For now, Im adding support only the common json() and text() and plan to add support for the remaining methods in a follow-up PR. If you think it's too large to review, I can limit it to parse() and move asResponse() to a separate PR.

@IlyasShabi
IlyasShabi requested a review from BridgeAR August 3, 2026 12:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants