PFIZER — BA, AI Testing · ROUND 4 FRONT · Elnaz Alipour, PhD (Solo) · Interviewer Intel + Opening + Technical Walkthroughs
Solo technical validator panel — Elnaz Alipour PhD (Sr Dir, Medical Analytics Care Gaps & Customer Segmentation) · She heard NONE of the Round 2 stories · All material is fresh · Go technical first
Target posture: peer-level quantitative rigor + pharma data credibility + GxP instinct
Framing: validated AI in regulated environments — doctoral-level precision is her native language
Elnaz Alipour, PhD — full intelligence profile
Who she is: Sr Director, Medical Analytics Care Gaps & Customer Segmentation, Pfizer (Feb 2024–present). PhD Physics/Biophysics, Brown University. 12+ yrs data science. Prior: Head of Data Science, Commercial Analytics, Veeva Systems (4 yrs).
Her career arc: Academic biophysicist (Northwestern, UConn postdocs) → Scotiabank (regulatory data, Hadoop/SQL) → Veeva (4 yrs: built propensity models, HCP prioritization, customer segmentation on Amazon Redshift with IQVIA/Symphony data) → Pfizer Sr Director. Pattern: progressively higher commercial and leadership scope. Data scientist who became a business leader — she speaks both.
Her world at Pfizer: Defines requirements for care gap models — identifies patients who should be on a medication but aren't. Builds HCP prioritization models — which physicians to target for commercial engagement. Customer segmentation. She is the consumer of AI analytics. The BA testing role validates the tools that serve her team's outputs directly.
Her public quote (Norstella): "Helps us understand barriers to treatments in multiple phases of the patient journey, examining disease signs, prescriptions, and reported outcomes." This is the lens she evaluates everything through — patient barriers and HCP behavior at the data level.
Ethical AI lens: Organizing committee, TMLS Women+AI Micro Summit. Vocal about fairness and ethical AI. If you probe for bias in segmentation models — that resonates with her personal values, not just her technical standards.
Her evaluation question: "Will this tester find the errors that corrupt my care gap outputs before my team acts on them?" Answer that question in everything you say.
Opening pitch — technical validator adaptation (≤60 sec)
"I'm a pharmacist by training — 21 years under FDA, NRC, and DEA — and an AI practitioner for the last decade. My background sits at the intersection that matters here: I've built regulated AI systems, and I've also been responsible for the clinical and compliance consequences when they fail."
Technical anchor (for Elnaz): "At NewsRx, I built a 4-blinded LLM evaluation harness solo — HHEM and BERTScore evaluation framework, 27 prompt iterations, 18% factual grounding improvement, 83% human preference rate. That's not a chatbot assessment — that's systematic model evaluation with documented evidence."
Pharma anchor: "At Amgen, I led validation of the first GxP-validated GenAI system for FDA regulatory submissions — 400+ user stories, 40% drafting time reduction. Board-approved expansion. Compliance instinct isn't something I learned here — it's embedded from two decades in regulated environments."
Bridge to her world: "The specific domain I'm adding depth to is commercial pharma analytics — the care gap and HCP targeting side. My SQL validation work at OptumInsight on Medicaid and TPA claims in Snowflake is the closest structural parallel. The methodology is the same; the schema is different."
Walkthrough — "How would you test an AI system you've never seen before?"
This is the highest-probability Elnaz question. She wants systematic thinking, not improvisation.
1Understand what it's supposed to do. Before touching the system: business stakeholder interview — what decisions does this AI support? What does correct look like? What have users complained about?
2Map the data access. What can the AI actually see? What does the user expect it to see? The gap between those two is where most failures live — before any prompting starts.
3Build a failure mode taxonomy first. Ask engineering: what failure modes have you already observed? Don't write test cases before you know what failure looks like in this specific system.
4Establish ground truth test set. Identify 10–20 questions I already know the correct answer to from source data. Run those first. The delta is the baseline.
5Systematize from the pattern. Hallucinations cluster — they don't occur randomly. Find the clustering rule: is it numerics? Proper nouns? Data freshness? Multi-step reasoning? Then stress-test that vector specifically.
For Elnaz specifically: "If I was testing an AI that generates care gap recommendations, my first ground truth set would be patients I can manually confirm have a care gap in the source data. Then I check whether the AI identifies them correctly — and more importantly, whether it fabricates gap records for patients who shouldn't be on the list."
Walkthrough — "Walk me through your prompt engineering process"
1Define expected output first. Write what the correct response looks like before writing the prompt. Prompt engineering without a target is optimization without a metric.
2Isolate variables. Change one thing at a time — system instruction, user template, context ordering, verbatim constraint. At NewsRx: 27 iterations, each change tracked, before/after measured. Otherwise you can't know what moved the needle.
3Constrain the failure mode. If the AI hallucinates numerics, add explicit verbatim constraints: "If a number is not in the source data, do not include it." If it misattributes — add attribution format requirements.
4Test the fix, not the prompt. Regression test: does the new prompt still pass the cases the old prompt handled? Prompt improvements that fix one thing and break three others are not improvements.
NewsRx anchor: "Hallucinations clustered on numerical specifics — drug dosages, publication dates. Root cause: under-constrained prompt + source metadata gaps. I added verbatim constraints for numerics and fixed the metadata schema. 18% grounding improvement, measured with HHEM and BERTScore. Every iteration was documented."
Walkthrough — "Describe your UAT process for AI"
1AI UAT differs from traditional UAT. Traditional: output is deterministic — pass/fail is binary. AI: output is probabilistic — acceptance criteria must account for output variance while still being measurable.
2Business acceptance criteria before engineering. What does the business stakeholder accept as "good enough"? Write the criteria jointly — not as an IT requirement. At Optinosis: SaMD UAT with CTO-approved acceptance criteria before a single test ran.
3Three-tier test suite. (a) Regression: must-pass known-good scenarios. (b) Edge cases: system limits and failure modes. (c) User acceptance: shadow sessions with actual end users doing real work.
4Document everything audit-ready. In pharma, "tested and passed" is not a deliverable. The evidence package — test cases, actual outputs, expected outputs, delta analysis, sign-off — that's the deliverable. At Amgen: GxP-validated. Every artifact traceable.
High-probability technical Q&A — Elnaz will go deeper
QuestionAnswer trigger — technical frame OK for Elnaz
"Tell me about a hallucination you identified and fixed" NewsRx anchor. Hallucinations clustered on numerical specifics — drug dosages, dates — not random. HHEM flagged factual drift; BERTScore showed semantic divergence. Root cause: under-constrained prompt + metadata gaps. Added verbatim constraints + fixed source schema. 18% grounding ↑, 83% human preference rate. 27 iterations documented. "Hallucinations cluster — they're diagnostic, not random noise."
"How do you distinguish prompt failure from data quality failure?" Same AI output, two different root causes. Write the ground truth SQL. If the source data is correct but the AI answer is wrong → prompt or retrieval failure. If the SQL also returns wrong data → data quality failure. Critical: fixing a prompt to compensate for bad data is hiding a data problem. Flag it upstream. At Pfizer: that's Jay Shetty's pipeline — not a prompt fix.
"How would you test for bias in a patient segmentation model?" Slice the test set by demographic subgroup — age, geography, race (where available). Check if model accuracy and recall are consistent across segments. If care gap identification rate diverges significantly by subgroup, that's a fairness signal. In pharma segmentation: under-identification of a demographic group could create a care access disparity — that's the clinical consequence. I'd flag that to Elnaz's team as a model validation finding, not just a testing artifact.
"What's your experience with Snowflake Cortex?" "SQL validation of AI outputs against Medicaid and TPA claims data at OptumInsight — used Snowflake for ground truth queries. Cortex modules conceptually: Analyst (NL→SQL, test the generated SQL not just the answer), Search (vector retrieval, test by comparing to SQL-returned ground truth), Complete (LLM inference, test for hallucinated facts against source records). The Cortex Agents layer I've studied depth for this role — multi-step reasoning traces are the hardest to test because failures are emergent across the chain."
"How do you handle a system you've improved that regresses after a data update?" Regression test suite should run on every data refresh, not just after prompt changes. If a data update breaks previously passing tests, the failure is in the ingestion pipeline — either a schema change, a source data quality drop, or a data freshness lag. Root-cause the pipeline change before touching the prompt. At NewsRx: daily ingestion meant daily regression exposure. The harness ran nightly.
"Why this role — you seem more senior?" "I want to embed in a pharma AI team where the testing problem is real and the stakes are high. Pfizer is building AI on commercial data that informs HCP targeting and care gap decisions — errors there have downstream clinical and commercial consequences. That's not a stepping stone — that's the specific intersection where my background actually matters. I've been on the builder side. I want to be on the validation side at scale."
Veeva bridge — use this with Elnaz specifically
She spent 4 years at Veeva. She built propensity scoring, HCP prioritization, and customer segmentation models on the Veeva Nitro platform (Amazon Redshift + IQVIA/Symphony Health data). She knows exactly what vendor-side AI analytics looks like when it's built right — and wrong.
Say: "I've been building the kind of AI systems your Veeva team was evaluating and deploying. The difference is I've also been on the side that has to defend model outputs to FDA and compliance — which means I test differently than someone who hasn't been accountable for regulated failures."
PFIZER — BA, AI Testing · ROUND 4 BACK · Business Q&A · Weak Spots · Questions for Elnaz · Watch-outs · Say This / Not That
Elnaz Alipour PhD — care gap + HCP prioritization lens · She will probe deeper than Rachael · Technical rigor is the currency · 90 sec answers · let data breathe
Round 4 posture: she's here to validate technical fit, not re-assess seniority
All Round 2 anchor stories are FRESH — she heard none of them
Why this role matters in Elnaz's world
Core thesis for Round 4: Elnaz's care gap models and HCP prioritization outputs are only as good as the AI tools feeding them. If the AI hallucinates a care gap or misclassifies an HCP, her team acts on wrong intelligence — with downstream commercial and patient consequences. This role validates that those tools are trustworthy before they reach her team.
The specific bridge: "What I'd be testing are the Cortex-based agents and chatbots that sit between your data lake and your analysts. If Cortex Analyst generates a subtly wrong SQL query on HCP prescribing data, the care gap model downstream gets corrupted inputs — and the error is invisible to the end user. That's the failure mode I'm built to catch."
HEOR connection: Elnaz listed HEOR as a top skill. HEOR models inform drug value and reimbursement decisions — AI is now #1 trend in HEOR (ISPOR 2026). If her team uses AI to accelerate outcomes research literature review or model development, she'll want a tester who understands that the AI output feeds a regulatory evidence package, not just a dashboard.
Business Q&A — Elnaz's commercial pharma lens
QuestionAnswer trigger
"What does good look like in 90 days?" Day 1–30: failure mode taxonomy for the highest-stakes AI system. Built from: business stakeholder interviews (Elnaz first), engineering orientation, user shadow sessions. Day 30–60: structured test plan executed, findings documented, root causes root-caused not surface-documented. Day 60–90: at least one measurable improvement validated. Deliverable: the team trusts the testing outputs — and has audit-ready evidence they can show upstream.
"How do you explain AI failures to business users?" "The AI got the wrong answer because it was looking at the wrong data, or because it misread the data it had." Two sentences. Then show the delta: here's what the AI said, here's what the source says. Excel format — AI output column, SQL ground truth column, delta flagged. No jargon. Anyone can read it and action it.
"How would you prioritize which system to test first?" Highest stakes first: which system is closest to a regulated or high-visibility decision? A Cortex Analyst tool generating care gap recommendations for HCP targeting is a higher priority than a general FAQ chatbot — the error consequence is commercial and potentially clinical. Secondary: frequency of user-reported complaints. I build the priority matrix collaboratively with business stakeholders — they know where the real risk lives. Signal to Elnaz: her care gap system is first in line.
"What's your pharma commercial data experience?" OptumInsight: SQL validation of GenAI outputs against Medicaid and TPA claims in Snowflake — HCP prescribing pattern data. Plus 21 years as nuclear pharmacist: I understand HCP prescribing from the dispensing side — why a physician orders what they order, what formulary and prior auth look like from the provider perspective. Commercial analytics is the dimension I'm adding depth to here. [Pause. Don't over-extend.]
Numbers — say these with precision
ProjectAnchor numbers
NewsRx27 iterations · 18% factual grounding ↑ · 83% human preference · 950+ doc throughput · 4-blinded harness · HHEM + BERTScore · hallucinations clustered on numerics
AmgenFirst GxP-validated GenAI at Amgen · 400+ user stories · 40% drafting ↓ · CTD Modules 3,4,5 · Board-approved expansion · FDA regulatory submissions
OptinosisSaMD UAT sign-off · SOC2/HIPAA · CTO-approved acceptance criteria · MCP agentic pipeline · audit-ready artifacts · on schedule
Invistics96% ML accuracy · 300+ hospitals · NIH SBIR $2.1M · Epic/Cerner/AllScripts/Meditech
OptumInsightSQL validation of GenAI outputs · Medicaid/TPA claims · Snowflake · HCP prescribing data
Weak spot defense — Elnaz will probe deeper than Rachael
GapHonest scopeWhat to say
Cortex hands-on thin OptumInsight: SQL validation against Snowflake. No direct Cortex hands-on. "I've used Snowflake for ground truth validation of AI outputs. Cortex modules conceptually: Analyst, Search, Complete, Document AI, Agents — I've studied depth for this role. The validation methodology doesn't change with the platform syntax. Cortex-specific schema and tooling I'd get current on week one."
Commercial pharma analytics domain Amgen is regulatory R&D, not HCP targeting or care gap analytics. "My pharma depth is primarily on the regulated clinical and regulatory side — FDA submissions, GxP validation, SaMD UAT. Commercial analytics is where I'm adding domain depth. OptumInsight is the structural bridge: claims-level HCP prescribing data, SQL-validated AI outputs in Snowflake."
"Why not an architect role?" Resume reads as AI Solutions Architect — overqualification risk more likely with Elnaz's seniority. "I want to be at the intersection where AI meets clinical consequence in pharma, at the depth of a team actually building and deploying these systems. I've been the architect. I want to be the person who makes sure the architecture actually works in production — under compliance, with documented proof."
Bias/fairness testing not prominent in resume Not explicitly listed but directly relevant to Elnaz's patient segmentation work. "Bias testing in segmentation models — I'd slice by demographic subgroup and check recall parity. If care gap identification rate diverges significantly by population subgroup, that's a fairness signal that has patient access implications. That's an area I'd want to build into the test framework from day one."
Chatbot interface QA vs. pipeline evaluation Background is eval harness / multi-agent systems, not chatbot-specific. "A chatbot is an agent with a conversational interface — same failure modes: hallucination, instruction adherence, edge case handling, refusal behavior. The methodology I've used scales directly. I've tested the harder version — multi-step agentic pipelines where failures are emergent across the reasoning chain."
Questions to ask Elnaz — signal depth to a PhD data scientist
  • Care gap data flow: "For the care gap models — are the AI-generated outputs going directly into the HCP prioritization scores, or is there a human review layer before recommendations reach the field? The testing priority shifts significantly depending on that answer."
  • Cortex Agents scope: "Is the AI layer primarily Cortex Agents orchestrating across structured and unstructured data, or are these simpler Cortex Analyst implementations? Agent testing is fundamentally different — failures are emergent across the reasoning chain."
  • Failure taxonomy: "Has the team already documented the known failure modes for these systems, or is building that taxonomy part of what this role would own from day one?"
  • Bias and fairness: "For the patient segmentation models — has there been any analysis of whether care gap identification rates are consistent across demographic subgroups? I'd want to understand the existing fairness baseline."
  • Validation maturity: "Has the validation approach for AI outputs in these workflows been adapted for probabilistic outputs specifically, or is it still closer to traditional IT validation with binary pass/fail criteria?"
  • Collaboration model: "When a testing finding requires a data pipeline change — is that a conversation between me and your data engineering team, or does it flow through a product owner first?"
Say this / not that — PhD technical validator version
Say this
  • "HHEM + BERTScore to measure grounding" (she understands)
  • "18% factual grounding improvement, 27 iterations"
  • "Hallucinations cluster — they're diagnostic, not random"
  • "Ground truth SQL vs. AI output — the delta is the finding"
  • "Bias testing: subgroup recall parity"
  • "Cortex Analyst: validate the generated SQL, not just the answer"
  • "Emergent failures across the agent reasoning chain"
  • "Probabilistic outputs need different acceptance criteria"
  • "GxP-validated — audit trail is the deliverable"
  • "Data quality failure — this is not a prompt fix"
Not that
  • "The AI makes mistakes sometimes" (too vague)
  • "vLLM on RunPod" (infrastructure noise)
  • "I'd need to learn the domain first" (confirms gap fear)
  • "I'm more of an architect than a tester" (undermines fit)
  • "I haven't worked with Cortex specifically" (stop there)
  • "Random hallucinations" (shows you don't know AI behavior)
  • "I'll defer to the data team on that" (loses Elnaz)
  • "I'm really looking to grow into commercial pharma" (confirms the fear)
  • "Business users can read the delta in Excel" (right for Rachael, too basic for Elnaz)
Best single framing line for Elnaz: "I build the evidence that lets a pharma organization trust its AI outputs — systematic, documented, and defensible under audit. That's the same rigor GxP validation demands, applied to probabilistic systems."
Watch-outs — Round 4 specific
  • She's technical — don't over-simplify. Unlike Rachael, Elnaz can handle BERTScore, HHEM, propensity model language. Calibrate to her level — she'll pull you deeper if she wants it. Don't dumb it down unprompted.
  • Round 4 = higher overqualification scrutiny. She's seen Rounds 1–3 pass. She may probe why you'd take a BA-level role. Answer: "this is the intersection, not a stepping stone." One line. Move on.
  • Don't conflate HCP prioritization with prescriber data basics. She built these models at Veeva. If you over-explain what HCP targeting is, she'll know you don't have depth. Say "prescribing propensity models" not "figuring out which doctors to call."
  • Bias/fairness is a values signal, not just a technical check. She's on TMLS Women+AI committee. If fairness in segmentation comes up — don't treat it as a box to check. Engage it as a substantive design concern in patient-facing AI.
  • Care gap bridge requires precision. Don't say "I understand care gaps." Say: "barriers to treatment across the patient journey — diagnosis, prescribing, adherence." Quote her own framing back.
  • 90-sec answers. Let data breathe. After an anchor number — 18%, 400+, 83% — pause. She's processing. Don't rush past your own evidence.
  • She may not ask much about Rachael's world. Rachael already sold. Elnaz is here to validate specific technical depth. Stay in that lane — don't drift into business communication framing mid-answer.
Credential chips — calibrated for Elnaz
RPh · 21yr nuclear pharmacist FDA · NRC · DEA GxP-validated AI SaMD UAT HHEM + BERTScore eval 4-blinded eval harness Multi-agent pipeline testing SQL ground truth validation Snowflake · HCP claims data Propensity model validation Regulated AI bias testing Sole contributor · no playbook Board-approved AI expansion NIH SBIR $2.1M ML system
Round 4 reminder: Elnaz heard NOTHING from Round 2. The NewsRx hallucination anchor, the Amgen GxP story, the lookup-vs-generation framing — all fresh. Use them.