When the number is wrong, the AI should say so

When the number is wrong, the AI should say so

Work you can act on and defend.

Work you can act on and defend.

Work you can act on and defend.

A P&T access and formulary analysis for a global pharmaceutical client, run in SteepRock Embedded Intelligence in August 2026. Every finding below is an actual output of that session.
A P&T access and formulary analysis for a global pharmaceutical client, run in SteepRock Embedded Intelligence in August 2026. Every finding below is an actual output of that session.
A P&T access and formulary analysis for a global pharmaceutical client, run in SteepRock Embedded Intelligence in August 2026. Every finding below is an actual output of that session.
September 2026
September 2026

Most AI does not say so

By June 2026, a public database had documented 1,598 court cases worldwide in which AI-fabricated content, including invented citations, fake quotes, and rulings that never existed, was filed with a court, and the tool involved was almost always a consumer chatbot the professional already had [1]. In the founding case, a lawyer confronted about fake citations asked the chatbot to confirm they were real, and it confirmed its own fabrications [2]. The pattern is not confined to law. An audit of 2.5 million biomedical papers, published in The Lancet in May 2026, found fabricated references in roughly 2,800 of them, at a rate that has risen twelvefold since 2023 [3]. In a global survey of clinicians across 15 specialties, 91.8 percent reported encountering AI hallucinations, and 84.7 percent believed those errors could cause patient harm [4]. Companies including Pfizer, Novartis, Sanofi, GSK, and AstraZeneca, five of the largest in the industry, have moved generative AI from pilots to enterprise scale in medical affairs, and their medical affairs teams must now demonstrate the accuracy and sourcing of AI-assisted work before it is used [6]. Three of those five companies use SteepRock Market Development Services, including Embedded Intelligence, and we see that requirement emerging everywhere. Every chatbot answer arrives fluent and confident, whether it is right or wrong. What follows is what it looks like when an AI checks its own number, fails it, and says so.

The number that did not add up

In August 2026, a global pharmaceutical company preparing for a potential launch asked SteepRock to run a pharmacy and therapeutics (P&T) access analysis across three major US health systems: a large academic medical center, an integrated payer-provider system, and a community health system. The client brought its own prompts for the work. The analysis ran in Embedded Intelligence (EI), connected to the client's engagement records in OLMS and to OLA's licensed data.

Partway through, the analysis produced a payer mix for one of the largest treatment programs in the country. The figure looked fine. It was built from 231 pharmacy claims.

The licensed national dataset in the same session held 447,957 claims for the same drug codes over the same window. One of the biggest programs in the country accounting for 0.052 percent of national volume is not plausible, and EI said so. It compared the institutional figure to the national total, failed it, declined to present the payer mix, and explained why: the connected linkage data tied only 160 providers of any specialty to the institution, and only 16 in the relevant specialty. The math was fine. The provider roster underneath it was thin. The fix is a better roster, not a better prompt.

A stand-alone assistant would have presented that payer mix. It has no national total to check against, so it reports what it finds, and it reports it with the same confidence whether the number is right or wrong.

What else the session found

Four more findings, each from data the client's own analysis had not reached.

A conflict the client had not priced in. Open Payments records showed the client had paid the head of the academic center's clinical program roughly $27,000 in consulting and related fees from 2023 to 2025, most of it in 2025. Under a standard conflict-of-interest policy, the most influential clinical voice on this formulary decision may be the one person required to recuse from it.

An assumption that ran backwards. Trial registry data showed both academic systems had active trial relationships with a competitor developing a product in the same class, and neither had a confirmed trial relationship with the client. The client's earlier analysis assumed the reverse.

A quality measure that penalizes new products at launch. The relevant quality measures run off coded value sets that will not include a newly approved product on day one. Patients started on it will score as non-compliant, automatically, until those value sets update. That argues for timing protocol changes to a calendar-year boundary.

Two corrected premises. EI flagged two assumptions in the client's own analysis as unsupported by the connected data and showed the records behind each. The client is reviewing both.

It also declined to do several things it was asked to do. It would not fabricate committee rosters, would not invent base rates, would not assert trial participation it could not confirm, labeled its reconstructions as reconstructions, and stopped to ask when a prompt was ambiguous.

Why it worked

Three things, and the model is not one of them. EI runs Claude, the same model available in any enterprise seat. What changes the result is what the model is connected to, the rules it works under, and how the session is run.

The data. EI ran in the client's own private, access-controlled workspace. Everything in the table below was queried live during the session. The plausibility check that caught the bad number exists only because the national total was in the room to compare against.

Embedded Intelligence

Data source

Value @ $144/hr

Stand-alone assistant

Your engagement records in OLMS

Connected

No access

~0.5

Licensed national pharmacy claims

Connected

No access

Provider-to-institution linkage

Connected

No access

Open Payments records

Connected

No access

Clinical trial registries

Connected

No access

Certified treatment center list (about 8,700 centers)

Connected

No access

Your approved product and context documents

Connected

No access

~2.2

The rules. Every question in EI, whether it comes from the team's shared prompt library or is typed fresh, runs under written rules that SteepRock engineers, tests, and maintains: no guessing where data is missing, every figure labeled as confirmed or inferred, gaps declared, never filled, and every claim traced to a connected source, an attached file, or the user's own input. Record-level claims link back to the original entry in OLMS. Public information is cited.

The working method. The client's own prompts for this use case, four of them, were well written by any standard. SteepRock's team still rebuilt them into a ten-prompt library before the session, making fourteen corrections, writing down the expected result for each prompt, and grading every output against that standard. The rebuild is part of why the bad number was caught: the client's original sequence asked for the payer breakdown first and the data-quality check second, and a caveat delivered after a number does not undo the number. The rebuild made the capture check a precondition. EI's rules would have labeled the 231-claim figure as unreliable; the rebuilt prompt let EI withhold it altogether. That is the difference between a platform and a platform with a tested prompt library, and SteepRock delivers both. The client's original prompts, and what the rebuild changed in each, are in the appendix.

Why this matters now

Medical affairs teams at large manufacturers are moving generative AI from pilots into everyday work, and the regulatory expectation is moving with them. FDA's January 2025 draft guidance and the FDA and EMA's joint Guiding Principles of Good AI Practice in Drug Development, published in January 2026, expect AI whose outputs can be audited, and the EU AI Act's high-risk obligations, deferred by the Digital Omnibus adopted in July 2026, apply from December 2, 2027 for standalone systems and August 2, 2028 for AI embedded in regulated products [5]. The practical requirement is the same in every case: a team has to be able to show where a number came from and why it can be trusted. A screenshot of a chat does not meet that bar. A documented session, with expected results written down before execution and every output reviewed against them, does.\

What this means for your team

Medical and scientific teams get answers grounded in real records, every claim labeled as confirmed or inferred, and limitations stated next to the findings.

Account and access teams get the questions that decide a launch answered from licensed data a general assistant cannot reach: who decides, who pays, who influences, and where competitors already stand.

Compliance teams get a paper trail: expected results documented before execution, every output reviewed against them, and every claim traceable to its source, delivered as part of the engagement.

Bring us your prompts

SteepRock works exclusively with life sciences manufacturers in pharma, biotech, and medical devices, and EI's rules were written for that regulated environment. If your team already has a set of prompts it relies on, send us your ten best. We will rebuild them, write down the expected results, run them against your connected data in EI, and show you what changes. The client's identity in this case study is protected; if you would like to hear their experience directly, contact us and we will ask.


Appendix: the prompts, as the client wrote them

The client arrived with four prompts for this use case. They were well written, and the rebuild still produced fourteen corrections across the ten-prompt library that ran. The four are shown here in the client's own words, anonymized only where names appeared. The rebuilt versions are not shown.

Prompt 3.1: Institutional access profile

“Build a P&T preparation profile for [Health System A]. Cover: the relevant patient volume flowing through it, the payer channel mix for that population, current prescribing concentration in the drug class and which prescribers drive it, the referral catchment feeding the institution, and the specialty decision-makers affiliated with it. Be explicit about which figures come from claims and which are inferred.”

What the rebuild changed. One prompt asking five questions returns five shallow answers. The rebuild separated it into three prompts covering decision structure, the institutional profile, and the influence network, and bound the unit of analysis, because a satellite facility and a formulary authority are not interchangeable.

Prompt 3.2: Payer channel and rejection lens

“For the relevant patient population over the last twelve months, break down pharmacy claims by payer channel and by paid, rejected and reversed status. Identify where rejections cluster and what that suggests about the access barriers we should expect to be raised in a P&T review.”

Prompt 3.3: Bias and capture check

“Before I read anything into that breakdown, tell me where this data is incomplete. For each payer channel, state whether it is reliably captured, partially captured or substantially under-captured in this dataset, and explain the mechanism behind any under-capture. Then tell me which of the conclusions I might reasonably have drawn from the previous output are not safe to draw, and restate the findings that do survive. Do not soften this — if a number should not be used, say so.”

What the rebuild changed. The best instinct in the set: the client knew their data had gaps and asked to have them exposed. The flaw was the order. By the time this prompt runs, the payer numbers are already on screen, and a caveat delivered after a number does not undo the number. The rebuild merged the two with the order reversed, making the capture check a precondition. That reversal is what let EI withhold the misleading figure.

Prompt 3.4: Public P&T pre-read

“Pull together what is publicly known about how large integrated health systems have structured recent P&T reviews for novel therapies in this class — evidence they weight, comparators they demand, and cost-effectiveness arguments that have landed. Turn it into an anticipated-questions pre-read with our supporting evidence mapped against each.”

What the rebuild changed. Replaced by a synthesis pre-read built from the outputs of the full ten-prompt library, so anticipated questions are grounded in the specific institutions under review. The rebuild also added six subjects the original set did not touch, covering order-set reality, institutional economics, competitive timing, research relationships, quality metrics, and discharge continuity, each of which produced one of the findings above.


Sources: [1] D. Charlotin, AI Hallucination Cases database, June 2026. [2] Mata v. Avianca, Inc. (S.D.N.Y. 2023). [3] Topaz M, Roguin N, Gupta P, Zhang Z, Peltonen L-M. Fabricated citations: an audit across 2.5 million biomedical papers. Lancet. 2026;407:1779-1781. [4] Medical Hallucinations in Foundation Models and Their Impact on Healthcare, MIT Media Lab and collaborators, arXiv:2503.05777 (2025); global clinician survey across 15 specialties. [5] FDA, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025; FDA and EMA, Guiding Principles of Good AI Practice in Drug Development, January 14, 2026; Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal of the EU, July 24, 2026. [6] IntuitionLabs, GenAI in Medical Affairs: Use Cases and Compliance Guardrails, revised April 2026. Client relationships per SteepRock records.

Appendix: the prompts, as the client wrote them

The client arrived with four prompts for this use case. They were well written, and the rebuild still produced fourteen corrections across the ten-prompt library that ran. The four are shown here in the client's own words, anonymized only where names appeared. The rebuilt versions are not shown.

Prompt 3.1: Institutional access profile

“Build a P&T preparation profile for [Health System A]. Cover: the relevant patient volume flowing through it, the payer channel mix for that population, current prescribing concentration in the drug class and which prescribers drive it, the referral catchment feeding the institution, and the specialty decision-makers affiliated with it. Be explicit about which figures come from claims and which are inferred.”

What the rebuild changed. One prompt asking five questions returns five shallow answers. The rebuild separated it into three prompts covering decision structure, the institutional profile, and the influence network, and bound the unit of analysis, because a satellite facility and a formulary authority are not interchangeable.

Prompt 3.2: Payer channel and rejection lens

“For the relevant patient population over the last twelve months, break down pharmacy claims by payer channel and by paid, rejected and reversed status. Identify where rejections cluster and what that suggests about the access barriers we should expect to be raised in a P&T review.”

Prompt 3.3: Bias and capture check

“Before I read anything into that breakdown, tell me where this data is incomplete. For each payer channel, state whether it is reliably captured, partially captured or substantially under-captured in this dataset, and explain the mechanism behind any under-capture. Then tell me which of the conclusions I might reasonably have drawn from the previous output are not safe to draw, and restate the findings that do survive. Do not soften this — if a number should not be used, say so.”

What the rebuild changed. The best instinct in the set: the client knew their data had gaps and asked to have them exposed. The flaw was the order. By the time this prompt runs, the payer numbers are already on screen, and a caveat delivered after a number does not undo the number. The rebuild merged the two with the order reversed, making the capture check a precondition. That reversal is what let EI withhold the misleading figure.

Prompt 3.4: Public P&T pre-read

“Pull together what is publicly known about how large integrated health systems have structured recent P&T reviews for novel therapies in this class — evidence they weight, comparators they demand, and cost-effectiveness arguments that have landed. Turn it into an anticipated-questions pre-read with our supporting evidence mapped against each.”

What the rebuild changed. Replaced by a synthesis pre-read built from the outputs of the full ten-prompt library, so anticipated questions are grounded in the specific institutions under review. The rebuild also added six subjects the original set did not touch, covering order-set reality, institutional economics, competitive timing, research relationships, quality metrics, and discharge continuity, each of which produced one of the findings above.


Sources: [1] D. Charlotin, AI Hallucination Cases database, June 2026. [2] Mata v. Avianca, Inc. (S.D.N.Y. 2023). [3] Topaz M, Roguin N, Gupta P, Zhang Z, Peltonen L-M. Fabricated citations: an audit across 2.5 million biomedical papers. Lancet. 2026;407:1779-1781. [4] Medical Hallucinations in Foundation Models and Their Impact on Healthcare, MIT Media Lab and collaborators, arXiv:2503.05777 (2025); global clinician survey across 15 specialties. [5] FDA, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025; FDA and EMA, Guiding Principles of Good AI Practice in Drug Development, January 14, 2026; Regulation (EU) 2024/1689 as amended by Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal of the EU, July 24, 2026. [6] IntuitionLabs, GenAI in Medical Affairs: Use Cases and Compliance Guardrails, revised April 2026. Client relationships per SteepRock records.

Phone

Want to speak with us directly?

Enter your phone number and we will give you a call

We can help you achieve your goals

For more than 20 years, SteepRock has served as a recognized thought leader and best in class strategic partner across the pharmaceutical, biotech, medical device, animal health, and nutrition industry segments. Your success is our success. We deliver technology, information and analytics to help support the most critical business decisions shaping the healthcare landscape and support the entirety of your business with AI making you and your team more efficient and responsive.

Copyright © 2025 SteepRock Inc. SteepRock is a registered trademark of SteepRock, Inc. All rights reserved.

Phone

Want to speak with us directly?

Enter your phone number and we will give you a call

We can help you achieve your goals

For more than 20 years, SteepRock has served as a recognized thought leader and best in class strategic partner across the pharmaceutical, biotech, medical device, animal health, and nutrition industry segments. Your success is our success. We deliver technology, information and analytics to help support the most critical business decisions shaping the healthcare landscape and support the entirety of your business with AI making you and your team more efficient and responsive.

Copyright © 2025 SteepRock Inc. SteepRock is a registered trademark of SteepRock, Inc. All rights reserved.

Phone

Want to speak with us directly?

Enter your phone number and we will give you a call

We can help you achieve your goals

For more than 20 years, SteepRock has served as a recognized thought leader and best in class strategic partner across the pharmaceutical, biotech, medical device, animal health, and nutrition industry segments. Your success is our success. We deliver technology, information and analytics to help support the most critical business decisions shaping the healthcare landscape and support the entirety of your business with AI making you and your team more efficient and responsive.

Copyright © 2025 SteepRock Inc. SteepRock is a registered trademark of SteepRock, Inc. All rights reserved.