PFIZER — BA AI Testing · ROUND 4 MASTER · Elnaz Alipour, PhD · April 15, 2026 · FRONT
Decision maker · Solo · ~30 min · Technical validator · ~80% decided — confirm fit, don't re-audition · She heard NONE of Round 2 stories — all anchors are fresh
Posture: peer-level quantitative rigor + pharma data credibility + GxP instinct
Her evaluation question: "Will this tester find the errors that corrupt my care gap outputs before my team acts on them?"
Elnaz Alipour PhD — full profile
Critical new intel: She was Head of Data Science, Commercial Analytics at Veeva Systems for 4 years (2020–2024) — built propensity models, HCP prioritization, and customer segmentation ON the Veeva Nitro platform using IQVIA/Symphony data. She knows Veeva from the inside. Surface-level Veeva talk = instant credibility loss.
At Pfizer now: Sr Director, Medical Analytics, Care Gaps & Customer Segmentation (Feb 2024–present). PhD Physics/Biophysics, Brown. 12+ yrs data science. Moved from Veeva (vendor) to Pfizer (buyer) — crossed over to close the gap Veeva couldn't. She defines what "correct" looks like for the AI tools she tests.
Her public lens: "Helps understand barriers to treatments across multiple phases of the patient journey — disease signs, prescriptions, reported outcomes." (Norstella) That is her filter on everything you say.
Ethics signal: TMLS Women+AI organizing committee. Vocal about fairness in AI. If bias/segmentation comes up — engage it as substantive, not a checkbox.
The Veeva pivot opportunity: At Veeva she optimized commercial engagement (HCP reach/frequency). At Pfizer she's optimizing clinical outcomes (right patients identified, treated). The gap between those two is exactly where AI testing matters most.
Opening pitch — ≤60 sec · technical validator frame
"I'm a pharmacist by training — 21 years under FDA, NRC, and DEA — and an AI practitioner for the last decade. My background sits at the intersection that matters here: I've built regulated AI systems, and I've been responsible for the clinical and compliance consequences when they fail."
Technical anchor: "At NewsRx: 4-blinded LLM evaluation harness, solo. HHEM + BERTScore. 27 iterations, 18% factual grounding improvement, 83% human preference rate. That's systematic model evaluation — not a chatbot assessment."
Pharma anchor: "At Amgen: first GxP-validated GenAI for FDA regulatory submissions. 400+ user stories, 40% drafting time reduction. Board-approved expansion. The compliance instinct is embedded — two decades of regulated environments."
"The domain I'm adding depth in is commercial pharma analytics — the care gap and HCP targeting side. My SQL validation work on Medicaid and TPA claims in Snowflake is the closest structural parallel. The methodology is identical. The schema is different."
30-min strategy
0–2Don't re-pitch. "I have a clear picture of what you're building from Rachael's conversation." Establish you're here to confirm fit, not audition.
2–20Her questions. Answer directly — no preamble. Every answer ends with bridge to: care gaps, HCP targeting, or segmentation model integrity. She's a physicist — go technical. She won't punish depth.
20–25Your 2 questions max. Lead: "What does the analytics function look like when the AI testing is working correctly — what changes about how your team operates?" Then governance if time.
25–30Close: "What I'd bring immediately is the ability to protect the integrity of the outputs your team depends on — documented evidence that holds up in a regulated environment. These outputs ultimately feed HCP and patient decisions. That's the testing role I want."
Walkthrough — "Test an AI you've never seen before"
Highest-probability question. She wants systematic thinking, not improvisation.
1Understand what it's supposed to do. Business stakeholder interview — what decisions does this support? What does correct look like? What have users complained about?
2Map the data access gap. What can the AI actually see vs. what the user expects it to see? That gap is where most failures live — before any prompting starts.
3Build failure mode taxonomy first. Ask engineering: what failure modes have you already observed? Don't write test cases before you know what failure looks like in this specific system.
4Establish ground truth test set. 10–20 questions I already know the correct answer to from source data. Run those first. The delta is the baseline.
5Systematize from the pattern. Hallucinations cluster — not random. Find the clustering rule: numerics? Proper nouns? Data freshness? Multi-step reasoning? Stress-test that vector.
For Elnaz: "For a care gap AI, my first ground truth set would be patients I can manually confirm have a care gap in source SQL. Then I check whether the AI identifies them — and whether it fabricates gaps for patients who shouldn't be on the list."
Walkthrough — "Prompt engineering process"
1Define expected output first. Write what correct looks like before writing the prompt. Optimization without a target is not engineering.
2Isolate variables. Change one thing at a time. NewsRx: 27 iterations, each tracked. Otherwise you can't know what moved the needle.
3Constrain the failure mode. Hallucinations clustered on numerical specifics at NewsRx — drug dosages, dates. Root cause: under-constrained prompt + metadata gaps. Added verbatim constraints. 18% grounding improvement.
4Test the fix, not just the prompt. Regression test: does the new prompt still pass the cases the old prompt handled?
Walkthrough — "UAT process for AI"
1AI UAT ≠ traditional UAT. Traditional: binary pass/fail. AI: probabilistic — acceptance criteria must account for output variance while remaining measurable.
2Business acceptance criteria before engineering. Write jointly — not as an IT requirement. At Optinosis: CTO-approved criteria before a single test ran.
3Three-tier suite. (a) Regression: must-pass known-good. (b) Edge cases: failure modes. (c) User acceptance: shadow sessions with actual end users doing real work.
4Audit-ready evidence package. At Amgen: GxP-validated. Test cases, actual outputs, expected outputs, delta analysis, sign-off. That's the deliverable — not "tested and passed."
Numbers — say these with precision
ProjectAnchors
NewsRx27 iterations · 18% grounding ↑ · 83% human preference · 4-blinded harness · HHEM + BERTScore · hallucinations clustered on numerics
AmgenFirst GxP GenAI for FDA · 400+ user stories · 40% drafting ↓ · CTD Modules 3,4,5 · Board-approved expansion
Invistics96% ML accuracy · 300+ hospitals · NIH SBIR $2.1M · Epic/Cerner/AllScripts/Meditech
OptumInsightSQL validation of GenAI outputs · Medicaid/TPA claims · Snowflake · HCP prescribing data
J&J MedTech$750K annual cost ↓ · zero production defects 2 yrs · FDA + EU MDR dual compliance
Cardinal HealthMDM $2M+/yr savings · $300M contract recovery · 80+ facility M&A integration
High-probability technical Q&A — she will go deeper
QuestionAnswer trigger
"Hallucination you identified and fixed" NewsRx anchor. Clustered on numerical specifics — not random. HHEM flagged factual drift, BERTScore showed semantic divergence. Root cause: under-constrained prompt + source metadata gaps. Added verbatim constraints + fixed schema. 18% grounding ↑, 83% human preference. 27 iterations documented. "Hallucinations cluster — they're diagnostic."
"Prompt failure vs. data quality failure" Write ground truth SQL. Source data correct but AI wrong → prompt or retrieval failure. SQL also wrong → data quality failure. Critical: fixing a prompt to compensate for bad data is hiding a data problem. Flag it upstream — at Pfizer, that's Shetty's pipeline, not a prompt fix.
"Test for bias in segmentation" Slice test set by demographic subgroup. Check whether care gap identification rate and recall are consistent across segments. Divergence = fairness signal with patient access implications. Flag to Elnaz's team as a model validation finding — not just a testing artifact. This hits her TMLS ethics work directly.
"Snowflake Cortex experience" "SQL validation against Medicaid/TPA claims at OptumInsight — Snowflake as ground truth. Cortex: Analyst (validate generated SQL, not just the answer), Search (vector retrieval vs. SQL ground truth), Complete (LLM inference, test for hallucinated facts), Agents (hardest — failures emergent across multi-step reasoning chain). Cortex-specific schema I'd get current on week one."
"Why this role — you seem more senior" "I want to be at the intersection where AI meets clinical consequence in pharma, at depth, inside a team actually building these systems. I've been the architect. I want to be the person who makes sure the architecture works in production — under compliance, with documented proof. The BA title is the entry point into the right org. I'm optimizing for domain, not title."
"What does good look like in 90 days?" Day 1–30: failure mode taxonomy for the highest-stakes system — built from stakeholder interviews (Elnaz first), engineering orientation, shadow sessions. Day 30–60: structured test plan executed, root causes root-caused not surface-documented. Day 60–90: one measurable improvement validated. Deliverable: team trusts testing outputs, has audit-ready evidence to show upstream.
Veeva — she built it. Match her level.
Do NOT describe Veeva as a CRM. She built the propensity scoring and HCP segmentation models on Veeva Nitro (Amazon Redshift + IQVIA/Symphony data) for 4 years. She knows what it delivers and what it doesn't.
The layers she knows: CRM = commercial engagement signals (reach, frequency, channel). Veeva Link = HCP identity + affiliation. Veeva Nitro/OpenData = prescriber analytics + IQVIA scripts. Veeva Compass = patient/prescriber longitudinal claims. Care gap logic lives in the patient data layer — upstream of the CRM layer.
"The Veeva layer gives you HCP engagement signals. The care gap logic sits upstream, in patient-level claims. The testing challenge is that the AI tools touching care gap identification are making inferences about clinical unmet need — not commercial activity. A CRM error misses a call. A care gap model error misses a patient's medication."
Orthogonal data — the data completeness bias
Definition: Independent, non-overlapping sources each capturing a different dimension. Elnaz's stack: claims (Rx/medical), lab/diagnostic, EMR/EHR, specialty pharmacy, patient-reported outcomes. Complete picture requires all five — but coverage is NOT uniform across populations.
The testing vector: Commercially-insured patients have dense EMR + claims coverage. Medicaid populations often have sparser specialty + lab coverage. A care gap model trained on dense coverage will systematically undercount gaps in sparse-coverage populations — even if the AI logic is perfect. That's data completeness bias, invisible unless you test by subgroup.
"I'd stratify care gap identification rate by data coverage tier — patients with all five sources vs. those with only two. Differential error rates across coverage tiers reveal model dependence on data completeness, not AI quality."
Data semantics — the silent failure mode
Semantic drift: Same term defined differently across systems. Example: "uncontrolled hypertension" = SBP > 140 in claims-derived care gap rule; AI fine-tuned on literature where "uncontrolled" = SBP > 130. Model identifies larger patient population than intended. Commercial team acts on inflated list. No one failed — the definitions diverged.
Semantic consistency test: Extract canonical definition from approved data dictionary or SQL rule. Test whether AI produces the same patient population as the SQL-defined rule. Patient in AI output but not SQL → semantic overreach. Patient in SQL but not AI output → semantic underreach. Both are semantic failures, not hallucinations.
"Semantic consistency testing is a separate test vector from hallucination testing. The AI can return a factually grounded response that still violates the business definition — it used 'care gap' the way a clinical paper defined it, not the way Pfizer's operational ruleset defines it."
Questions to ask Elnaz — 2 max
Q1 (Vision): "What does the analytics function look like when the AI testing is working correctly — what changes about how your team operates?" Shows you're thinking at her level, not the task level.
Q2 (Governance): "For the care gap models — are AI outputs going directly into HCP prioritization scores, or is there a human review layer before recommendations reach the field? The testing priority shifts significantly depending on that answer."
Backup if time: "Has there been any analysis of care gap identification rate consistency across demographic subgroups? I'd want to understand the existing fairness baseline."
PFIZER — BA AI Testing · ROUND 4 MASTER · BACK · 4-Quadrant Identity · Weak Spots · Say-This · Watch-Outs
Use the 4-quadrant frame when asked "What makes you different?" · Every cross-domain story ends anchored to: care gaps, HCP targeting, or segmentation model integrity
The anchor rule: "The methodology transfers. The domain changes. What stays constant:
what happens to the patient when the AI is wrong?"
The 4-Quadrant Identity — the unique positioning
The claim: Most candidates have one quadrant — maybe two. Bishop activates all four simultaneously. This is not a generalist weakness. It is the rarest architecture in healthcare AI.
1/4 CLINICAL Registered Nuclear Pharmacist + RPh (1984–present). 21 yrs under NRC/FDA/DEA. Compounding, QC, patient dosing, clinical consequence of failure — learned from patients receiving contaminated isotope doses at age 25.
1/4 BUSINESS OPS MBA (Rollins/Crummer). LSSBB. P&L ownership ($23M regional). JIT supply chain national scale. MDM at Cardinal Health ($2M+ savings, $300M contract recovery). Built $870M acquisition.
1/4 TECHNICAL DARPA pattern recognition 1979 → production ML (Invistics 96%, NIH, peer-reviewed) → GenAI (Amgen, first GxP FDA system) → LLM eval harness (NewsRx, HHEM+BERTScore) → EHR integration (Epic, Cerner, AllScripts, Meditech, Snowflake, Databricks, AWS).
1/4 REGULATORY NRC + FDA + DOT + EPA + OSHA + State Boards simultaneously 20 years. Then: HIPAA, GxP, ICH Q-series, EU MDR, EUDAMED, GDPR, CMS, DEA. Not policy familiarity — operational accountability with patient-safety consequences.
LSSBB Thread across all four: DMADV (Syncor service line, $6M+), DMAIC (Cardinal MDM), 8D root cause (1984 contamination crisis), SAFe 6.0 zero-rework (Amgen 400+ user stories), zero-defect GxP delivery (J&J 2 yrs).
"Most consultants have one of these quadrants. I have all four — applied simultaneously across nuclear medicine, provider, payer, biopharma, and medtech. When a care gap model produces an output, I can evaluate it clinically, technically, from a business consequence standpoint, and from a regulatory defensibility standpoint — in the same conversation."
Best all-4-quadrant story — Invistics drug diversion (use if asked "What makes you different?" or "Complex situation?")
Clinical Diversion = patient safety crisis. Assembled expert panel: physicians, nurses, pharmacists, hospital admins, law enforcement, regulators. I spoke every language in that room.
Ops Deployed to 300+ facilities. NIH SBIR $2.1M to Wolters Kluwer acquisition. Investigation time 4–20 hrs → 10–30 min.
Technical Canonical data model across 8+ EHR vendors + 3+ ADC vendors. Single ML model, 300+ facilities. 96.3% accuracy. Named contributor, AJHP peer-reviewed.
Regulatory HIPAA + GDPR + DEA-relevant. Output = risk categories, not binary accusations. Human-in-loop mandatory — false positives have employment and legal consequences.
"The reason I built risk tiers — not binary flags — is the same reason I'd test a care gap model: the consequence of a false positive is not a data error. It's a physician targeted incorrectly, or a patient who doesn't get treatment. I design tests around what happens when the AI is wrong, not just when it's right."
Weak spot defense — she will probe deeper than Rachael
GapHonest scopeSay this
Cortex hands-on thin OptumInsight: SQL validation in Snowflake. No direct Cortex hands-on. "I've used Snowflake for ground truth validation of AI outputs. Cortex conceptually — Analyst, Search, Complete, Document AI, Agents. Validation methodology doesn't change with platform syntax. Cortex-specific schema I'd get current on week one."
Commercial pharma analytics domain Amgen is regulatory R&D, not HCP targeting or care gap analytics. "My pharma depth is primarily on the regulated clinical and regulatory side — FDA submissions, GxP, SaMD UAT. Commercial analytics is where I'm adding domain depth. OptumInsight is the structural bridge: claims-level HCP prescribing data, SQL-validated AI outputs in Snowflake."
"Why not architect?" Resume reads senior. Overqualification risk with Elnaz's seniority. "The BA title is the entry point into pharma commercial AI testing at the right org. I've been the architect. I want to be the person who makes sure the architecture actually works in production — under compliance, with documented proof. I'm optimizing for domain, not title."
Bias/fairness not prominent Not listed on resume but directly relevant to her segmentation work and TMLS ethics role. "Bias testing in segmentation models: slice by demographic subgroup, check recall parity. If care gap identification rate diverges significantly by population — that's a patient access disparity, not just a testing artifact. I'd build that into the framework from day one."
Chatbot QA vs. pipeline Background is eval harness / multi-agent systems, not chatbot-specific. "A chatbot is an agent with a conversational interface — same failure modes: hallucination, instruction adherence, edge case handling. I've tested the harder version — multi-step agentic pipelines where failures are emergent across the reasoning chain."
Business Q&A — compressed
QuestionAnswer trigger
"How do you explain AI failures to business users?" "The AI got the wrong answer because it was looking at the wrong data, or because it misread the data it had." Two sentences. Then show the delta: AI output column, SQL ground truth column, delta flagged. No jargon. (Don't dumb this down for Elnaz — she can read the SQL herself. Show her the delta directly.)
"Prioritize which system to test first?" Highest stakes first — closest to regulated or high-visibility decision. Cortex tool generating care gap recommendations for HCP targeting is higher priority than a general FAQ chatbot. Secondary: frequency of user-reported complaints. Build priority matrix collaboratively with business stakeholders. Signal: Elnaz's care gap system is first in line.
"Pharma commercial data experience?" OptumInsight: SQL validation of GenAI outputs against Medicaid/TPA claims in Snowflake — HCP prescribing pattern data. Plus 21 years nuclear pharmacist: understand HCP prescribing from the dispensing side — why a physician orders what they order, formulary and prior auth from the provider perspective. Commercial analytics = the layer I'm adding here. [Pause. Don't over-extend.]
Credential chips — calibrated for a PhD data scientist
RPh · 21yr nuclear pharmacist FDA · NRC · DEA GxP-validated AI SaMD UAT HHEM + BERTScore eval 4-blinded eval harness Multi-agent pipeline testing SQL ground truth validation Snowflake · HCP claims data Regulated AI bias testing Board-approved AI expansion NIH SBIR $2.1M ML system Sole contributor · no playbook AJHP peer-reviewed contributor
Say this / not that — PhD technical validator version
Veeva is a CRM system pharma companies use to track HCPs
→
The Veeva CRM layer captures commercial engagement signals. The care gap logic sits in patient-level claims data — upstream and a different accountability level.
I'd test the AI output against the expected answer
→
I'd extract the canonical SQL definition of each term from your data dictionary and validate whether the AI produces the same patient population. Delta = semantic failure, not hallucination.
I'd test across different types of patients
→
I'd stratify care gap identification rate by data coverage tier. Differential error rates across coverage tiers reveal model dependence on data completeness — not AI quality.
I haven't worked with Cortex specifically
→
Cortex-specific schema I'd get current on week one. The validation methodology — ground truth SQL vs. AI output, semantic consistency testing, emergent failure detection in agent chains — transfers directly.
Random hallucinations
→
Hallucinations cluster — they're diagnostic, not random noise. Find the clustering rule; stress-test that vector specifically.
I'm a fast learner and can pick up whatever tools you use
→
The Snowflake Cortex stack — Analyst, Search, Agents — is the exact environment I've been preparing for. Cortex Agents are the complex case: failures are emergent across the reasoning chain. I'm ready for that.
Business users can read the delta in Excel
→
I'd show her the delta directly — AI output vs. SQL ground truth. She can read the SQL herself. The evidence package is audit-ready, not just readable.
Best single framing line: "I build the evidence that lets a pharma organization trust its AI outputs — systematic, documented, defensible under audit. That's the same rigor GxP validation demands, applied to probabilistic systems."
Watch-outs — Round 4 specific
  • She's technical — don't over-simplify. She can handle BERTScore, HHEM, propensity model language. Don't dumb it down unprompted.
  • Don't describe Veeva at surface level. She built it. Match her level or bridge past it. Never say "CRM system."
  • Don't over-explain HCP targeting. She built propensity models on it for 4 years. Say "prescribing propensity models" — not "figuring out which doctors to call."
  • Bias/fairness is a values signal. She's on TMLS Women+AI committee. Engage it as a substantive design concern — not a box to check.
  • Care gap bridge requires precision. Don't say "I understand care gaps." Say: "barriers to treatment across the patient journey — diagnosis, prescribing, adherence." Her own framing.
  • Round 4 = higher overqualification scrutiny. She's seen Rounds 1–3 pass. Answer in one sentence: "The BA title is the entry point into the right org. I'm optimizing for domain, not title." Move on.
  • 90-sec answers. Let data breathe. After an anchor number — 18%, 400+, 96% — pause. She's processing. Don't rush past your own evidence.
  • Stay in the technical lane. She's not Rachael. Don't drift into business communication framing mid-answer. She already bought the business credibility.
The round context — why you're here
RoundWhoTheir worldResult
R2John Pastor
Dir, Business Technology
Global Data Mgmt, system integration, data pipelines. Not analytics — plumbing.IT credibility. 50 min + 12 min overage. ✓ Passed.
R3Rachael Rathbun + Shetty
(MedConnect 148 sources, $35M)
Commercial analytics delivery + the data lake infrastructure that feeds Elnaz's models.Cultural fit. They endorsed you TO Elnaz. ✓ Passed.
R4Elnaz Alipour PhDCare gaps, HCP prioritization, patient segmentation. She is the decision maker.~80% decided. Confirm technical judgment fit.
All Round 2 stories are FRESH. NewsRx hallucination anchor, Amgen GxP, Invistics ML system, the semantic/orthogonal framing — Elnaz heard none of it. Use everything.