PFIZER — BA, AI Testing · ROUND 2 FRONT · Interviewer Intel + Conceptual Walkthroughs + Weak Spot Defense
Business stakeholder panel — Elnaz Alipour PhD (Sr Dir, Medical Analytics) + Rachael Rathbun · 2026-04-09 4:00 PM ET · The Judge Group / Mark Haber
Target posture: pharma-credible AI practitioner who can operate independently
Framing: regulated AI + measurable outcomes + business value — NOT vLLM or infrastructure
Interviewer intel — Elnaz Alipour, PhD
Profile: Sr Director, Medical Analytics Care Gaps & Customer Segmentation at Pfizer. PhD Physics/Biophysics (Brown). 12 years data science. Prior: Head of Data Science, Commercial Analytics at Veeva Systems (4 yrs).
Her world: HCP prioritization, care gap models, prescribing analytics, customer segmentation — Snowflake-native. She evaluates AI through commercial pharma data quality and ROI lenses.
What she's listening for:
  • Does candidate understand pharma commercial data? (HCP targeting, claims, segmentation)
  • Will they produce measurable, defensible improvement?
  • Can they operate without hand-holding — she's senior and busy.
Bridge to her world: "OptumInsight — I ran SQL validation of GenAI outputs against Medicaid/TPA claims in Snowflake. The AI classified HCP behavior; I confirmed against the underlying records. That's care gap work." OptumInsight is your commercial claims bridge. Use it.
Veeva angle: She spent 4 years as Head of Data Science at Veeva. She knows exactly what vendor-side AI looks like when it's built right — and when it isn't. John has been the person building it on the vendor side (Amgen, NewsRx, Optinosis). Say: "I've built the kind of AI she was evaluating at Veeva."
Interviewer intel — Rachael Rathbun
Role unknown — treat as business owner of an AI tool. She cares about reliability, business value, and whether you communicate clearly. She is the person who has to tell her team the AI tool can be trusted.
  • For Rachael: Plain language. Concrete outcomes. "Here's what I found, here's what I fixed, here's how we know it's better."
  • She will not ask about BERTScore. She will ask: "How do you explain an AI failure to someone who doesn't know AI?"
  • Best signal to send Rachael: "I produce documentation that business stakeholders can understand without a translator."
Opening pitch — business reframe (≤60 sec)
"I'm a pharmacist by training — 21 years under FDA, NRC, and DEA — and an AI practitioner for the last decade. I've been building and validating regulated AI systems at the intersection where errors have real consequences."
Then anchor: "At Amgen, I led validation of the first GxP-validated GenAI system for FDA regulatory submissions — 400 user stories, 40% drafting time reduction, board-approved for expansion." Pause. Let it land.
Then close: "At NewsRx, I built a 4-blind LLM evaluation harness solo — 27 prompt iterations, 18% factual grounding improvement. That's the combination I bring: I can find what's wrong, root-cause it, fix it, and show you the delta."
  • Do NOT mention: vLLM, RunPod, Qwen3-32B, infrastructure — too technical, wrong room.
  • DO use: "measurable improvement," "audit-ready," "pharma compliance," "business user trust."
Conceptual walkthrough — "Walk me through how you'd embed data for an AI system"
1Identify source. What should the AI know? Product data, SOPs, clinical guidelines, HCP data from Snowflake. Scope defines quality ceiling.
2Chunking. Break into pieces that preserve meaning. Too small = loses context. Too large = retrieval imprecision. Related content — same HCP, same product — should stay together.
3Embedding model. Converts text to vectors (numbers). Domain-tuned models understand pharma terminology. General-purpose models miss clinical nuance.
4Storage. Vector database. Snowflake Cortex Search has native vector support — no separate infrastructure needed in this environment.
5Test I run. Ask the AI a question I know the answer to from the source data. Hallucination usually means: (a) bad chunking, (b) wrong embedding model, or (c) data was never loaded. Diagnose systematically.
Business translation for Rachael: "If the AI doesn't know something it should, the data either wasn't put in, or it was put in incorrectly. I find which one."
Conceptual walkthrough — "Walk me through how you'd use Snowflake to validate an AI response"
1AI makes a claim → write SQL against the underlying table to check it. Pattern: AI classifies → SELECT confirms from source record.
2Three failure modes: (a) retrieval failure — AI didn't get the right records. (b) Prompt failure — got the right records, misinterpreted them. (c) Data quality — records were incomplete or wrong at the source.
3Cortex Analyst (NL→SQL): validate by checking generated SQL against expected query structure. Most failures are AI querying the wrong column — not wrong logic.
4Excel for stakeholder communication: AI output column, SQL result column, delta flagged. Rachael can read that. Engineering can action it.
Signal to Elnaz: "OptumInsight — I did exactly this: SQL validation of GenAI outputs against Medicaid/TPA claims in Snowflake. The pattern doesn't change; only the schema changes."
Conceptual walkthrough — "Walk me through who you'd go to when you join"
1Business stakeholders first. Understand what the AI is supposed to do and what's frustrating users. Listen before touching anything.
2Engineering team. What data does the AI have access to? How are prompts structured? What guardrails exist? What's the model?
3End users — shadow sessions. See real behavior before running formal tests. Users find edge cases faster than test plans.
4Data team. What's in Snowflake? Data freshness? What can the AI actually see vs. what users expect it to know?
5First artifact: failure mode taxonomy — not a test plan. You can't write good tests until you know what failure looks like in this specific system.
Conceptual walkthrough — "How would you prioritize which system to test first?"
Primary criterion: Highest stakes first. Which system is closest to a regulated or high-visibility workflow? A care gap model with AI-generated HCP recommendations — higher priority than a general FAQ chatbot.
Secondary signal: Frequency of user complaints. If business users have lost trust in a specific tool, testing it first restores confidence fastest.
Elnaz frame: "If the care gap model has AI-generated HCP targeting recommendations, that's the one I want to test first — that's where an error has downstream commercial consequence."
Weak spot defense table — know these cold
GapHonest scopeWhat to say
Snowflake depth thin OptumInsight — 3 months validation queries. No Cortex hands-on. "I've validated AI outputs against Snowflake source data — the pattern is AI produces, SQL confirms against source. Cortex module syntax I'd get current on in week one; the validation logic doesn't change."
Commercial pharma domain Amgen is regulatory R&D (CTD modules). Not commercial analytics / HCP targeting. "OptumInsight: Medicaid/TPA claims analytics. 21 years nuclear pharmacist — I understand HCP prescribing from the dispensing side. Commercial analytics is the dimension I'm adding depth to here."
"Why this role — you're more senior?" Resume is framed as AI Solutions Architect — overqualified read risk. "I want to embed in a pharma AI team where the testing problem is real and the stakes are high. This is the intersection I want to work at — not a stepping stone, the destination."
Chatbot QA vs. pipeline testing Background is eval harness / multi-agent pipelines, not chatbot interface QA specifically. "A chatbot is an agent with a conversational interface — same failure modes: hallucination, instruction following, edge cases, refusal behavior. The methodology is identical. I've tested the harder version."
Excel validation not on resume Not listed explicitly — could read as gap. "Yes — AI output in one column, source data in another, delta flagged. That's a standard tool in my validation workflow for stakeholder communication."
Key operating principles — drop these naturally
Fix upstream, not the output Failure taxonomy before test plan Expected outputs written before running tests Hallucinations cluster — test where data is thin Root cause → precise fix engineering can execute Audit-ready artifacts every time Probabilistic outputs need different acceptance criteria Business stakeholders first — before touching the system
Best identity line: "I make AI systems reliable enough to trust — in pharma, under compliance, with documented proof."
Numbers to know cold
ProjectAnchor numbers
NewsRx27 prompt iterations · 18% grounding ↑ · 83% human preference · 950+ doc throughput · 4-blinded harness · HHEM + BERTScore
AmgenFirst GxP-validated GenAI at Amgen · 400+ user stories · 40% drafting ↓ · CTD Modules 3,4,5 · Board-approved expansion
OptinosisSaMD UAT sign-off · SOC2/HIPAA · CTO acceptance criteria · MCP agentic · audit-ready docs · on schedule
OptumInsightSQL validation of GenAI outputs · Medicaid/TPA claims · Snowflake
Invistics96% ML accuracy · 300+ hospitals · NIH SBIR $2.1M · Epic/Cerner/AllScripts/Meditech
PFIZER — BA, AI Testing · ROUND 2 BACK · Business Panel Q&A · Positioning · Questions to Ask · Watch-outs
For Elnaz Alipour PhD (commercial analytics lens) + Rachael Rathbun (business owner lens) · Show: pharma depth + independent operation + business communication
Pattern: answer with anchor → bridge to their world → close with outcome metric
Never fill silence. 90 sec most answers. 2 min max for anchor stories.
What this role is in business terms (Elnaz + Rachael's world)
Core thesis: Pfizer's analytics teams are building AI on top of commercial data — HCP segmentation, care gaps, prescribing models. This role tests whether those systems produce outputs that business users can trust and act on.
Elnaz's frame: Her care gap models inform HCP targeting decisions. If the AI recommends the wrong HCP or misclassifies a prescribing pattern, there's downstream commercial consequence. Testing is risk management.
Rachael's frame: She owns a tool her team uses. If it gives wrong answers, her team loses time and loses trust. She wants someone who can diagnose and fix reliably — without requiring her to translate between business and engineering.
Business impact language to use:
  • Faster AI adoption — testing removes friction to deployment
  • Reduced hallucination risk — direct compliance and commercial consequence
  • Business user trust in AI outputs — Rachael's actual concern
  • Audit-ready validation — Elnaz's scientific credibility requirement
Why John for this specific team
Elnaz has Veeva background. She knows what good vendor-side AI looks like when it's built right — she evaluated it for 4 years. John has been the person building it. "I've built the kind of systems she was evaluating at Veeva."
Pharma-native + AI — rare combination. Most AI testing candidates know AI. John knows pharma AND AI. 21 years under FDA/NRC/DEA. Amgen GxP. Optinosis SaMD. The compliance instinct is embedded, not learned in prep.
Sole contributor capability. NewsRx: entire eval framework built alone, no playbook, no team. Can operate independently in unstructured environment — which the JD explicitly calls for and a Sr Director appreciates.
Quantified every engagement. Every project above has defensible numbers. Not "I improved the system" — "18% grounding improvement, 27 iterations, documented." That's the scientific rigor Elnaz (PhD Physics) respects.
Coaching — how to land in a business stakeholder panel
  • Lead with business outcome, not technical mechanism. "18% improvement in factual accuracy" lands. "HHEM score delta across blinded eval cohorts" does not.
  • Name the consequence. "If this care gap model misclassifies a prescriber, the rep targets the wrong HCP — that's a commercial cost." Show you understand why it matters.
  • Let Elnaz lead technically. She's a PhD data scientist. Don't over-explain. If she probes deeper, match her level. If she stays business, stay business.
  • Rachael signal: Speak directly to her — "and from your perspective, as the person whose team uses this tool, what that means is..." Acknowledge both audiences.
  • Pause after anchor numbers. 18%. 40%. 83%. Let them breathe. Don't rush past your own evidence.
  • Don't over-defend weak spots. Name the scope honestly, redirect to what you bring. One sentence on the gap, two sentences on the strength.
High-probability business stakeholder questions — answer triggers
QuestionAnswer triggers
"Why do you want this role?" "I want to embed in a pharma AI team where the testing problem is real and the stakes are high. Pfizer is building AI on commercial data that drives HCP targeting decisions — errors there have real consequences. That's the intersection I want to work at." Do not say "career transition" or "learning opportunity."
"How would you work with our team?" Business stakeholders first — understand what the AI should do and what's frustrating. Engineering second — understand the architecture and data access. My deliverable to both: a findings document that business can read and engineering can action. I'm the translator, not a gatekeeper.
"Tell me about a time you improved an AI system's reliability" NewsRx: hallucinations clustered on numerical specifics — not random noise. HHEM + BERTScore flagged the pattern. Root cause: under-constrained prompt + source metadata gaps. Fixed both. 18% grounding improvement, 83% human preference rate. Documented every step. "Hallucinations cluster — test where data is thin."
"How do you explain AI failures to business users?" "The AI got the wrong answer because it was looking at the wrong data, or because it misread the data it had." One of those two things is almost always true. Then show the delta — here's what it said, here's what the source says, here's how we fixed it. Excel format. No jargon.
"What does good look like in 90 days?" Day 1–30: failure mode taxonomy for the highest-stakes system, built from business stakeholder interviews + engineering orientation + user shadow sessions. Day 30–60: structured test plan executed, findings documented, root causes identified. Day 60–90: at least one measurable improvement implemented and validated. "Good" = the team trusts the testing outputs.
"How do you handle disagreement with engineering?" I bring verbatim inputs, outputs, and expected behavior — not opinions. Engineers fix precisely defined problems. My job is to define them precisely enough that the disagreement becomes a discussion about the data, not about interpretation. At Amgen, that framing was how we got complex refusal-behavior specs through review.
"What's your experience with pharma commercial data?" OptumInsight: SQL validation of GenAI outputs against Medicaid/TPA claims in Snowflake — HCP prescribing patterns, claims-level data. Plus 21 years as nuclear pharmacist — I understand HCP prescribing behavior from the dispensing side. My pharma depth is primarily regulatory; commercial analytics is the dimension I'd be adding depth to here. [Pause. Don't over-extend.]
"How do you prioritize across multiple AI systems?" Highest stakes first: which system is closest to a regulated or high-visibility decision? A care gap model informing HCP targeting is higher priority than a general FAQ tool. Frequency of user complaints is a secondary signal. I build the priority matrix collaboratively with business stakeholders — they know where the real risk lives.
Questions to ask — signal depth to Elnaz specifically
  • Care gap model architecture: "For the care gap and customer segmentation models — are AI-generated recommendations embedded directly in the outputs that reps act on, or is there a human review layer before they reach the field?"
  • Snowflake Cortex scope: "Is Cortex the primary AI layer on top of the warehouse, or are there separate model deployments I'd be testing against? Testing approach differs significantly."
  • Failure taxonomy maturity: "Has the team documented known failure modes for the current AI systems — what kinds of errors show up most frequently — or is building that taxonomy part of what this role would own?"
  • Collaboration model: "When a testing finding requires a prompt or data change — is that a joint session with engineering, or do I write a spec and they implement? I can work either way."
  • 90-day signal (for Rachael): "From your perspective — what's the thing that, if I fixed it in the first 90 days, would make the biggest difference to your team's confidence in the AI tools?"
  • Validation approach: "For AI outputs in compliance or analytics-adjacent workflows — has the validation approach been adapted for probabilistic outputs, or is it still closer to traditional IT validation?"
Watch-outs for this round
  • Don't go technical in a business panel. Elnaz is technical — she'll probe if she wants depth. Rachael is not. Stay at business outcome level unless Elnaz explicitly pulls you deeper.
  • Don't over-defend the commercial pharma gap. One sentence: scope your OptumInsight experience honestly. Two sentences: redirect to what you bring. Move on.
  • Don't lead with "I'm overqualified but..." The seniority question is a trap. Answer: "this is the intersection I want to work at — not a stepping stone."
  • Watch for the "tell me about yourself" reframe. In round 2 they already know your resume. They're probing fit and communication style. Lead with business framing, not career history.
  • 30-minute clock: 90 sec most answers. 2 min max for anchor stories. Brevity signals confidence. Don't fill silence — they're processing, not waiting.
  • Rachael silence: If she doesn't ask much, address her directly once: "From a day-to-day usability standpoint — are there specific behaviors users have flagged as unreliable?"
Say this / not that — business panel version
Say this
  • "18% improvement in factual accuracy"
  • "Business users can read the delta in Excel"
  • "Care gap model errors have commercial consequence"
  • "First artifact is a failure mode taxonomy"
  • "I've built the kind of systems Elnaz was evaluating at Veeva"
  • "Audit-ready documentation — not just a findings list"
  • "The test I run first is a question I already know the answer to"
Not that
  • "HHEM score delta across blinded cohorts" (Rachael lost you)
  • "vLLM deployment on RunPod" (wrong room)
  • "I haven't worked with Snowflake Cortex specifically" (stop there)
  • "I'm more of an architect than a tester" (undermines fit)
  • "The model hallucinates randomly" (shows you don't know AI)
  • "I'd need the engineering team to look at that" (loses Elnaz)
  • "I'm really looking to grow into commercial pharma" (confirms the fear)
Best single framing line for this panel: "I make AI systems reliable enough that business teams can trust them — in pharma, under compliance, with documented proof of improvement."
Credential chips — drop these when relevant
RPh · 21yr nuclear pharmacist MBA LSSBB GxP-validated SaMD UAT FDA submissions NRC / DEA Snowflake SQL HCP prescribing data Medicaid/TPA claims Sole contributor Board-approved AI Commercial analytics bridge