PFIZER ROUND 4 — CRITICAL STRATEGIC INTEL · Elnaz Is The Decision Maker · Round 4 Posture Shift + Cortex Agents Crash Course
New intelligence: Technical people → Rachael → both report to Elnaz · Her team endorsed you · She's doing final leader evaluation, not another tech screen · Posture shifts from "prove it" to "fit with her vision"
Supplement to Round 4 Cheat Sheet · Use alongside existing Round 2 Primer
Read this page first — strategic framing overrides tactical Q&A
★ DECISION STRUCTURE — READ BEFORE ANYTHING ELSE
Confirmed interview chain:
Round 1: Two technical people (engineering/data science) → PASSED
Round 2: Rachael Rathbun, AI Strategy Lead R&D (13yr Pfizer) → PASSED (stayed 12 min over, strong engagement)
Rounds 3+: [Additional rounds not in DB]
Round 4: Elnaz Alipour, PhD — Sr Director → Rachael and the technical people both report to her.
What "she needs to talk to me" means: Rachael and/or the technical team didn't just pass you — they explicitly recommended you up the chain to their boss. In a pharma org, that's a strong internal sponsor signal. The decision is not starting from zero at Round 4 — it's 70–80% made.
Posture shift: You are not here to prove you belong. You are here to confirm for Elnaz that what her team told her about you is true — and to show her you can work inside HER vision. Stop defending. Start collaborating.
She's evaluating three things only:
  • Can I trust this person's judgment in my org without constant oversight?
  • Will they protect the integrity of outputs my team depends on?
  • Do they understand what I'm actually trying to build?
What changed: Round 2 vs. Round 4 posture
Round 2 (Rachael)Round 4 (Elnaz)
Business owner of an AI toolSr Director who owns the analytics org
Evaluating: can you communicate clearly?Evaluating: can I trust your judgment?
Needed business outcome framing firstPhD physicist — lead with rigor
Prove you can work with engineeringProve you understand what HER team builds
No recommendation going inHer team explicitly recommended you
50/50 decision pendingDecision ~80% made — she's confirming fit
Stay business, don't go technicalGo technical when she pulls — she'll pull
Anchor stories untestedAll anchor stories FRESH — use them all
Opening — collaborative leader frame (not defend-and-prove)
"I've had the chance to speak with Rachael and the engineering team and I have a much clearer picture of what you're building. What I'd bring is the ability to protect the integrity of those analytical outputs — systematically, with documented evidence, at the rigor that a regulated pharma environment requires."
Then the Amgen anchor: "I led GxP-validated AI at Amgen — first of its kind there. 400+ user stories, FDA regulatory submissions. The compliance instinct that requires isn't something I'm learning — it's embedded."
Then the NewsRx anchor: "At NewsRx I built an evaluation harness solo — HHEM and BERTScore framework, 27 prompt iterations, 18% factual grounding improvement. That's not assessment — that's systematic model validation with evidence."
Close: "I understand the AI outputs that feed Rachael's team ultimately inform the care gap and HCP decisions your org depends on. Testing those tools isn't a support function — it's a data integrity function. That's the role I want."
★ CORTEX AGENTS — NEW IN 2025 · NOT IN ROUND 2 PRIMER · KNOW THIS COLD
What Cortex Agents are: Autonomous AI agents that orchestrate multi-step reasoning across Snowflake's structured and unstructured data. They combine Cortex Analyst (SQL), Cortex Search (semantic), and LLMs into a single reasoning chain — breaking complex questions into sub-tasks and executing them sequentially. Data never leaves Snowflake.
Why this is critical for Round 4: The JD says Pfizer built "several sophisticated AI agents." The BA role tests those agents. Cortex Agents are almost certainly the primary subject of this role — not simple chatbots. Testing agents is fundamentally harder than testing a single LLM call.
How Agents Work
User asks: "List patients with care gaps who saw an HCP in the last 90 days"
↓
Agent decomposes → Sub-task 1: get care gap patients (Cortex Analyst → SQL)
↓
Sub-task 2: find HCP visits in last 90 days (Cortex Analyst → SQL)
↓
Sub-task 3: cross-reference and synthesize (LLM reasoning)
↓
Agent returns composed answer with sources
Agent Failure Modes — harder than single-LLM testing
  • Emergent failures: Each sub-task may be individually correct but the composition produces a wrong answer. This is the hardest failure to catch — no individual step fails, but the chain logic is wrong.
  • Sub-task hallucination: Agent fabricates an intermediate result and uses it as input to the next step. Error compounds through the chain.
  • Incorrect tool selection: Agent uses Cortex Search when it should use Cortex Analyst — retrieves semantically similar documents instead of querying the structured table. Returns plausible-looking wrong data.
  • Reasoning loop or dead end: Agent can't complete a sub-task, infers a placeholder, and continues. Silent partial answer passed as complete.
How to Test Cortex Agents
  • Decompose the trace: For every agent run, inspect the full reasoning chain — what sub-tasks did it create, in what order, with what tool calls? The trace is the test artifact.
  • Unit test each sub-task: Can each step pass independently? If Sub-task 1 SQL is wrong, Sub-tasks 2–3 inherit corrupted inputs. Validate bottom-up.
  • Ground truth for the composition: Write the manual SQL that would answer the full question. Compare end result. Delta = agent composition failure, even if each step looked clean.
  • Tool selection testing: Ask questions that are ambiguous between structured lookup (Analyst) and semantic retrieval (Search). Verify agent chose the correct tool for the data type.
  • Regression suite for chains: Agent behavior is sensitive to prompt changes in any step. Full chain regression must run after any modification — not just the modified step.
Say this to Elnaz: "Testing Cortex Agents is fundamentally different from testing a single LLM call — failures are emergent across the reasoning chain. My approach is to decompose the agent trace into unit tests for each sub-task first, then validate the composition result against ground truth SQL for the full question. That catches both individual step failures and chain-level logic errors."
Cortex AI Observability — new GA 2025 · relevant to testing
What it is: Built-in evaluation and tracing for Cortex AI functions. Logs all AI function calls, model selections, reasoning steps, and outputs for audit and quality monitoring.
Testing relevance: Observability is the infrastructure that makes systematic AI QA possible inside Snowflake. For a BA tester, this means: (a) you can trace the exact reasoning path that produced a bad output, (b) you can compare output quality before and after a prompt change with an audit trail, (c) regulators can review AI decision logs — this is the GxP angle Elnaz will care about.
Say: "AI Observability in Cortex gives me the trace I need to root-cause which step in a multi-step agent chain produced the failure — and it generates the audit evidence for a regulated environment without additional tooling."
Cortex Code Governance Skills — data quality as a test vector
What it is: AI-powered governance in Snowflake Cortex Code. Includes: (a) Data Quality Skill — health scoring, anomaly detection, SLA alerting, root cause analysis for failing metrics, table comparison for migration validation. (b) Data Governance Skill — natural language answers about access control, audit trails, permissions, role hierarchies.
Testing relevance: When an AI output is wrong, the question is: was it the AI or the data? Cortex Code Governance gives you a native tool to check data quality health before declaring a prompt failure. If the upstream data quality score is degrading — the AI isn't wrong, the data is. This is a precise diagnostic capability, not a general-purpose tool.
Say to Elnaz: "My QA background directly maps to Cortex's data quality monitoring — before I diagnose a prompt failure, I check the data health score for the upstream table. A degrading quality metric is upstream of the AI, not a prompt fix. That distinction is the difference between fixing the right thing and papering over a pipeline problem."
Snowflake Horizon — governance + access tiers (Elnaz's world)
What it is: Snowflake's universal governance platform. Auto-classifies data (PII, PHI, confidential), applies dynamic masking policies, enforces row-level access. All Cortex AI functions respect Horizon access controls — the AI cannot retrieve data a user's role cannot see.
Access tier pattern (enterprise standard): Public → Internal → Confidential → Sensitive/PHI → Highly Restricted. In pharma: HCP-level PII and patient data sit in the highest tiers. Cortex outputs must not expose PHI that a given role cannot access — even through an AI-generated summary.
Testing angle for Elnaz's team: "I'd add data leakage testing to the suite — confirm that a Cortex response generated from restricted-tier data doesn't surface PHI for a user whose role should only see aggregated outputs. This is a governance test vector that most AI testing frameworks miss."
HEOR + Care Gap AI mapping — Elnaz's domain in Cortex terms
Her quote: "barriers to treatments in multiple phases of the patient journey — disease signs, prescriptions, outcomes."
Elnaz's domainCortex toolTesting angle
Care gap analytics dashboardCortex Analyst (NL→SQL)Validate generated SQL queries correct table/time range for gap calculation
HCP targeting recommendationsCortex Agents (multi-step)Decompose agent trace; validate each sub-task; compare final output to manual SQL
Medical literature / HEOR reviewCortex Search (semantic)Compare retrieved docs to SQL-queryable ground truth; check semantic accuracy not just keywords
Clinical document processingDocument AI (AI_EXTRACT)Compare extracted fields against source documents; NPI / date accuracy critical
Patient segment definitionCortex Analyst + HorizonValidate SQL logic; add subgroup bias testing (recall parity across demographics)
Data quality monitoringCortex Code Data QualityHealth score baseline; SLA alerting before AI outputs degrade
Key insight for Elnaz: The AI tools Bishop would be testing are the upstream inputs to her care gap and HCP prioritization models. A testing failure is not a QA metric — it's a data integrity problem that corrupts her team's analytics. Frame it that way.
Questions that signal strategic fit — ask Elnaz these
  • Vision: "What does the analytics function look like when the AI testing is working correctly — what changes about how your team operates?"
  • Priority: "Rachael mentioned [paraphrase what Rachael discussed] — are the care gap AI tools the first priority, or is there a different system with more immediate risk?"
  • Governance: "For outputs that feed into HCP targeting — what does the validation evidence package need to look like for your team to be able to act on AI recommendations with confidence?"
  • Bias/fairness: "Has there been any baseline fairness analysis on patient segmentation model outputs — whether care gap identification rates are consistent across demographic subgroups?"
  • Collaboration: "When I find a finding that requires a data pipeline change vs. a prompt change vs. a model change — what does escalation look like in your team's workflow?"