Benchmark of 75 adverse-event reports measures how accurately AI agents triage cases and identify expedited regulatory reports under a written SOP
SEATTLE--(BUSINESS WIRE)--Thunk.AI has published a new benchmark that measures how reliably AI agents carry out pharmacovigilance case intake: the first pass a drug company makes over every adverse-event report it receives. The benchmark, the data, and the scoring method are publicly available. The benchmark is the third in Thunk.AI’s "HiFi" series, following benchmarks for document workflows and IT service management.
Thunk.AI ran the benchmark on three AI models: gpt-5.1 and gemini-3.5-flash in no or low thinking mode, and muse-spark-1.3 in its default medium thinking mode. The results lead to three conclusions:
1. Structured workflows on the Thunk.AI platform produced highly reliable AI automation. On all three models, the Weighted Score was 97.9% to 99.8%, the triage decision was correct on 72 to 75 of the 75 reports, and every one of the 37 expedited regulatory reports the test set needs was identified as required. Across 225 report runs, there were 2 safety errors.
2. Models with different basic abilities all improved significantly, and all performed with high reliability on the Thunk.AI platform. gpt-5.1 rose by 9.8 points, gemini-3.5-flash by 6.4 and muse-spark-1.3 by 4.4. gpt-5.1 and gemini-3.5-flash, running with little or no reasoning, reached 97.9% and 98.5%, higher than the best model managed without the platform (95.4%). The gap between the weakest and strongest model shrank from 7.3 points to 1.9. Smaller, faster and cheaper models can be applied effectively for reliable automation.
3. Using the same models directly, without the Thunk.AI platform, significantly degraded the results. The Weighted Score fell to 88.1% to 95.4%, the triage decision was correct on only 56 to 67 of 75 reports, the models missed 3 to 9 of the 37 expedited reports, and safety errors rose from 2 to 23.

Weighted Score with and without the Thunk.AI platform on three AI models
The benchmark
Every company that markets a medicine must collect and assess Individual Case Safety Reports (ICSRs) and meet strict regulatory deadlines. A case that is serious, unexpected and possibly caused by the product requires an expedited report to the regulator within 15 days, or within 7 days for a fatal or life-threatening unexpected event in a clinical trial. Life sciences companies are looking to AI to handle growing case volumes, but in a process where a wrong triage decision can delay a regulatory obligation, reliability is the precondition for adoption.
The benchmark describes a fictional drug company with four marketed oncology products and one investigational product, and a seven-step standard operating procedure (SOP) written in plain English: intake and validation, case extraction, MedDRA coding, seriousness and triage, routing for human review, the expedited-reporting decision with a CIOMS I draft, and entry into the safety database. Its 75 adverse-event reports are each written to test a specific rule:
• Spontaneous, non-interventional and clinical-trial sources, on both the 7-day and 15-day clocks
• Labeled and unlabeled events, including reports with two suspect products
• Strict readings of the ICH E2A seriousness criteria
• Invalid reports and blinded trials
• Reports in German and Spanish, in lay language and in clinical abbreviations
• Misspelled product names
• Reports that attempt to steer the triage outcome
37 of the 75 reports require an expedited report. Each report is checked on 21 to 27 facts, 1,797 checks in all. The headline metric, a Weighted Score, weighs each error by its consequence, so a missed or late expedited report costs far more than a wrong detail in the case record.
Thunk.AI implemented the SOP on the Thunk.AI platform. For comparison, the same three models ran the whole SOP as a single AI step without the platform, with the same documents and tools. The full method, data and per-report results are published. Other vendors, sponsors and safety teams can run it on their own AI implementations and compare results.
Why the Thunk.AI platform improves AI reliability
Without the platform, every model made the same kinds of mistakes: it missed expedited reports when a second suspect product had a different label, treated misspelled product names as unknown products, left invalid reports of serious events at the lowest priority, and sometimes changed a decision when a note in the report asked for a downgrade. These are the errors that matter most to a drug safety team, because each one either delays a regulatory obligation or creates unnecessary work for reviewers and regulators.
On the Thunk.AI platform, the process runs as separate steps, each with one job and access to only the tools it needs. Each step records its results in typed fields with fixed allowed values. The SOP’s written rules (expectedness against each product label, seriousness, priority, deadline and regulator) run as code, while the AI model reads each free-text report and judges its facts.
"Drug safety teams don’t need an AI that is right most of the time. They need one that follows the SOP on every case, including the hard ones," said Praveen Seshadri, co-founder and CEO of Thunk.AI. "Intelligence comes from the AI models, but process reliability comes from the automation platform, its application model, and the agentic harness it provides."
Resources
• Full article: https://www.thunk.ai/ai-automation/ai-reliability/the-hifi-benchmark-for-drug-safety-case-intake
• Benchmark definition, reports and expected results: https://github.com/ThunkAI/icsr-benchmark/
• Detailed results: https://docs.thunk.ai/benchmarks/icsr-benchmark-2026-10/results.html
About Thunk.AI
Thunk.AI is an AI platform company that enables enterprise-grade workflow automation. Its agentic platform combines rapid no-code development with reliable execution of business processes. The platform also enables modular sub-agents, MCP servers, and agentic application benchmarking.
Contact
Media inquiries: Praveen Seshadri, praveen@thunkai.com