| Question | Answer trigger — technical frame OK for Elnaz |
|---|---|
| "Tell me about a hallucination you identified and fixed" | NewsRx anchor. Hallucinations clustered on numerical specifics — drug dosages, dates — not random. HHEM flagged factual drift; BERTScore showed semantic divergence. Root cause: under-constrained prompt + metadata gaps. Added verbatim constraints + fixed source schema. 18% grounding ↑, 83% human preference rate. 27 iterations documented. "Hallucinations cluster — they're diagnostic, not random noise." |
| "How do you distinguish prompt failure from data quality failure?" | Same AI output, two different root causes. Write the ground truth SQL. If the source data is correct but the AI answer is wrong → prompt or retrieval failure. If the SQL also returns wrong data → data quality failure. Critical: fixing a prompt to compensate for bad data is hiding a data problem. Flag it upstream. At Pfizer: that's Jay Shetty's pipeline — not a prompt fix. |
| "How would you test for bias in a patient segmentation model?" | Slice the test set by demographic subgroup — age, geography, race (where available). Check if model accuracy and recall are consistent across segments. If care gap identification rate diverges significantly by subgroup, that's a fairness signal. In pharma segmentation: under-identification of a demographic group could create a care access disparity — that's the clinical consequence. I'd flag that to Elnaz's team as a model validation finding, not just a testing artifact. |
| "What's your experience with Snowflake Cortex?" | "SQL validation of AI outputs against Medicaid and TPA claims data at OptumInsight — used Snowflake for ground truth queries. Cortex modules conceptually: Analyst (NL→SQL, test the generated SQL not just the answer), Search (vector retrieval, test by comparing to SQL-returned ground truth), Complete (LLM inference, test for hallucinated facts against source records). The Cortex Agents layer I've studied depth for this role — multi-step reasoning traces are the hardest to test because failures are emergent across the chain." |
| "How do you handle a system you've improved that regresses after a data update?" | Regression test suite should run on every data refresh, not just after prompt changes. If a data update breaks previously passing tests, the failure is in the ingestion pipeline — either a schema change, a source data quality drop, or a data freshness lag. Root-cause the pipeline change before touching the prompt. At NewsRx: daily ingestion meant daily regression exposure. The harness ran nightly. |
| "Why this role — you seem more senior?" | "I want to embed in a pharma AI team where the testing problem is real and the stakes are high. Pfizer is building AI on commercial data that informs HCP targeting and care gap decisions — errors there have downstream clinical and commercial consequences. That's not a stepping stone — that's the specific intersection where my background actually matters. I've been on the builder side. I want to be on the validation side at scale." |
| Question | Answer trigger |
|---|---|
| "What does good look like in 90 days?" | Day 1–30: failure mode taxonomy for the highest-stakes AI system. Built from: business stakeholder interviews (Elnaz first), engineering orientation, user shadow sessions. Day 30–60: structured test plan executed, findings documented, root causes root-caused not surface-documented. Day 60–90: at least one measurable improvement validated. Deliverable: the team trusts the testing outputs — and has audit-ready evidence they can show upstream. |
| "How do you explain AI failures to business users?" | "The AI got the wrong answer because it was looking at the wrong data, or because it misread the data it had." Two sentences. Then show the delta: here's what the AI said, here's what the source says. Excel format — AI output column, SQL ground truth column, delta flagged. No jargon. Anyone can read it and action it. |
| "How would you prioritize which system to test first?" | Highest stakes first: which system is closest to a regulated or high-visibility decision? A Cortex Analyst tool generating care gap recommendations for HCP targeting is a higher priority than a general FAQ chatbot — the error consequence is commercial and potentially clinical. Secondary: frequency of user-reported complaints. I build the priority matrix collaboratively with business stakeholders — they know where the real risk lives. Signal to Elnaz: her care gap system is first in line. |
| "What's your pharma commercial data experience?" | OptumInsight: SQL validation of GenAI outputs against Medicaid and TPA claims in Snowflake — HCP prescribing pattern data. Plus 21 years as nuclear pharmacist: I understand HCP prescribing from the dispensing side — why a physician orders what they order, what formulary and prior auth look like from the provider perspective. Commercial analytics is the dimension I'm adding depth to here. [Pause. Don't over-extend.] |
| Project | Anchor numbers |
|---|---|
| NewsRx | 27 iterations · 18% factual grounding ↑ · 83% human preference · 950+ doc throughput · 4-blinded harness · HHEM + BERTScore · hallucinations clustered on numerics |
| Amgen | First GxP-validated GenAI at Amgen · 400+ user stories · 40% drafting ↓ · CTD Modules 3,4,5 · Board-approved expansion · FDA regulatory submissions |
| Optinosis | SaMD UAT sign-off · SOC2/HIPAA · CTO-approved acceptance criteria · MCP agentic pipeline · audit-ready artifacts · on schedule |
| Invistics | 96% ML accuracy · 300+ hospitals · NIH SBIR $2.1M · Epic/Cerner/AllScripts/Meditech |
| OptumInsight | SQL validation of GenAI outputs · Medicaid/TPA claims · Snowflake · HCP prescribing data |
| Gap | Honest scope | What to say |
|---|---|---|
| Cortex hands-on thin | OptumInsight: SQL validation against Snowflake. No direct Cortex hands-on. | "I've used Snowflake for ground truth validation of AI outputs. Cortex modules conceptually: Analyst, Search, Complete, Document AI, Agents — I've studied depth for this role. The validation methodology doesn't change with the platform syntax. Cortex-specific schema and tooling I'd get current on week one." |
| Commercial pharma analytics domain | Amgen is regulatory R&D, not HCP targeting or care gap analytics. | "My pharma depth is primarily on the regulated clinical and regulatory side — FDA submissions, GxP validation, SaMD UAT. Commercial analytics is where I'm adding domain depth. OptumInsight is the structural bridge: claims-level HCP prescribing data, SQL-validated AI outputs in Snowflake." |
| "Why not an architect role?" | Resume reads as AI Solutions Architect — overqualification risk more likely with Elnaz's seniority. | "I want to be at the intersection where AI meets clinical consequence in pharma, at the depth of a team actually building and deploying these systems. I've been the architect. I want to be the person who makes sure the architecture actually works in production — under compliance, with documented proof." |
| Bias/fairness testing not prominent in resume | Not explicitly listed but directly relevant to Elnaz's patient segmentation work. | "Bias testing in segmentation models — I'd slice by demographic subgroup and check recall parity. If care gap identification rate diverges significantly by population subgroup, that's a fairness signal that has patient access implications. That's an area I'd want to build into the test framework from day one." |
| Chatbot interface QA vs. pipeline evaluation | Background is eval harness / multi-agent systems, not chatbot-specific. | "A chatbot is an agent with a conversational interface — same failure modes: hallucination, instruction adherence, edge case handling, refusal behavior. The methodology I've used scales directly. I've tested the harder version — multi-step agentic pipelines where failures are emergent across the reasoning chain." |