LEXISNEXIS / RELX — Senior ML Engineer · FRONT · Technical / ML Architecture
RAG pipeline design, multi-agent orchestration, hallucination mitigation, LLM evaluation, production ML systems
Product: Protege (LexisNexis flagship AI assistant — replaces Lexis+ AI Feb 2026)
Interviewers: Gaurav Singh (strategy) + Thiyagarajan Selvaraj (technical bar-raiser)
Opening pitch — say it exactly (≤60 sec)
Posture: Production ML engineer who operates in regulated environments where hallucination is a liability problem, not a quality score.
"I'm a pharmacist by training and an AI systems practitioner by the last decade. What makes me unusual: I've built and shipped multi-agent AI in regulated environments where citations must be verifiable — pharma FDA submissions, SaMD platforms, NIH-funded ML at 300+ hospitals.

At NewsRx I built self-hosted LLM infrastructure — vLLM, Qwen3-32B — and a four-blinded multi-agent evaluation harness as sole contributor. At Amgen: first GxP-validated GenAI for FDA regulatory submissions, multi-agent architecture, 400+ user stories, 40% drafting time reduction.

Protege lives at exactly this intersection — domain-critical AI that cannot get citations wrong. I know how to build the system and how to know when it's wrong."
RAG pipeline design — how to answer
Opening line: "Start with chunking strategy — it determines everything downstream."
Chunking: Semantic > fixed-size for legal docs. Preserve citation context — don't split case citations across chunks.
Retrieval: Hybrid BM25 + dense (maximize recall). Cross-encoder reranker before LLM sees results (boost precision).
Protege angle: "Your knowledge graph traversal post-retrieval for authority ranking is the right pattern — higher-court rulings over lower. I'd want to understand where the graph boundary sits in your retrieval pipeline."
Eval: RAGAS (faithfulness, answer relevancy, context recall). LLM-as-judge + human eval for novel failures.
RAG vs. fine-tuning decision framework
SituationUse
Knowledge changes frequently (case law)RAG
Need source traceability / citation verifiabilityRAG
Hallucination is liabilityRAG + verifier agent
Stable domain vocab, specific output formatFine-tune
Latency-sensitive, high-frequency taskDistill → small model
LexisNexis strategyBoth — teacher→student distillation for speed
High-probability Q&A — Thiyagarajan (technical deep-dive)
Multi-agent workflow reliability?
Orchestrator owns state; individual agents stateless + idempotent. Explicit failure contracts — structured result with confidence + failure reason. Reflection agent reviews before delivery (same pattern Protege uses). NewsRx: 4-blinded agents + enrichment agent cross-correlating — extend that here.
Design a verifiable citation system?
Generator and verifier must be separate — same model cannot reliably verify its own output. Citation grounding: every cited case runs through retrieval verification before response finalized. Force structured output (JSON with case_id, jurisdiction, authority_level). Fallback: explicit "could not verify" — never silently omit.
Hallucination in production?
Upstream: verbatim-only constraint for citations/numerics + explicit refusal spec when source unavailable. Architecture: Shepard's pattern — independent citation verification agent. Monitoring: HHEM + BERTScore as pipeline gates, not post-hoc. NewsRx lesson: hallucinations cluster at knowledge gaps — test where data is thin.
Handle agent loops / failures?
Max step budget per agent + explicit termination conditions. Agent returns structured ToolResult with success/failure/reason — orchestrator decides retry vs. escalate vs. fallback. Context window: summarize intermediate results before passing; don't accumulate raw traces.
LLM eval methodology?
Metric set: factual grounding (primary), BERTScore (semantic), HHEM (hallucination), citation accuracy, preference rate. 27-iteration controlled experiments, 1 variable per iteration, full audit trail. RAGAS for RAG-specific (faithfulness, answer relevancy, context recall).
Self-hosted vs. API LLMs?
Self-hosted: data sovereignty, no per-token cost at scale, full control, can fine-tune — GPU infra overhead. API: speed to market, model flexibility (swap GPT-5→Claude 4 without infra change). NewsRx: vLLM + Qwen3-32B on RunPod — eliminated API costs at 950+ doc throughput. LexisNexis hybrid is correct: proprietary legal data on-prem, general AI queries via API.
High-probability Q&A — Gaurav (strategy / stats)
Production ML end-to-end?
NewsRx: sole contributor, vLLM infra + 4-agent eval harness + PostgreSQL + Streamlit + GxP governance. 4-person staffing plan delivered alone in 8 weeks. Beyond contracted scope: DB, deployed web app, SME review system — became MVP engagement foundation. Quantified: 950+ doc throughput, 18% grounding ↑, 83% human preference, 27 iterations.
ML system drift / monitoring?
Feature drift (input distribution) vs. concept drift (label relationship) — different responses. MLflow for experiment tracking + model registry; shadow deployment before cutover. Invistics: 96% accuracy maintained across 300+ hospitals — EHR diversity was the stress test. Alert on output distribution change — embed quality metrics in inference path.
Domain you don't know?
Step 1: domain expert interviews — what does 'wrong' look like to a lawyer? Step 2: black-box exploration before reading architecture docs (surfaces failures unfiltered). Step 3: failure mode taxonomy. Amgen: told to figure out GxP-validated GenAI with no playbook. Same approach. Shipped.
Ambiguity in ML requirements?
Write expected outputs BEFORE running tests — or you'll rationalize failures as acceptable. Document assumptions explicitly in test plan. Stakeholder walkthrough BEFORE engineering gets the change list.
ML architecture response formulas
RAG pipeline
Chunk strategy → hybrid retrieval → rerank → LLM → verify → deliver
Agent reliability
Stateless agents → structured contracts → orchestrator state → reflection gate → deliver
Hallucination control
Constrain upstream → separate verifier → embed metrics as gates → audit trail
Production ML
Baseline → shadow deploy → MLflow tracking → drift monitor → cutover
Say this, not that
Do say
  • "Generator and verifier must be separate."
  • "I write expected outputs before running tests."
  • "Hallucinations cluster at knowledge gaps."
  • "I embed quality metrics in the inference path."
Do not say
  • "I'm not a deep coder."
  • "I'd use LangChain for that."
  • "Legal AI is different from what I've done."
  • "I haven't used Cortex specifically."
Technical chips
vLLMQwen3-32BRAG Graph RAGHHEMBERTScore RAGASMLflowRunPod multi-agentreflection agentdistillation hybrid retrievalcross-encoder reranker GxP-validatedcitation verifiability
LEXISNEXIS / RELX — Senior ML Engineer · BACK · Strategy / Interviewers / Differentiators
Interviewer profiles, anchor stories, differentiators, questions to ask, watch-outs, abbreviations
Interview: 2026-04-01 · 3pm–4pm EST · Via Ascendion / Rajat Patel
Job# 5106/RELNJP00032533 · 12mo + ext + CTH · 100% Remote
Interviewer intel — Gaurav Singh (hiring mgr)
  • Role: Director, Data Science · LexisNexis · 6 months in · Reports to Chief AI Officer Min Chen
  • Background: IIT Kanpur M.Stat → SAS risk/collections modeling → Tiger Analytics consulting director → LexisNexis
  • His lens: statistical rigor, production delivery, LLM fundamentals. Consulting background = values adaptability and ambiguity tolerance.
  • What he's building: new AI feature rollout for Protege, team ramping now for build phase. Funding aligned to CAO.
  • Growth mindset signal: William Standard noted he's "looking for people with growth mindset." Lead with learning velocity, not just credentials.
Interviewer intel — Thiyagarajan Selvaraj (tech bar-raiser)
  • Role: Platform/infrastructure engineer embedded in ML product team · LexisNexis hackathon winner
  • Background: Cognizant → LexisNexis. Strong .NET/Java + cloud (AWS, Azure, CloudFormation). Innovation Olympiad winner.
  • Described as: "Very sharp on everything ML" — brings to ALL interviews as consistent bar-raiser.
  • His lens: Can your ML code run in production? Containerization, latency, monitoring, clean Python. Systems thinking applied to ML.
  • Expect: RAG mechanics deep-dive, agent failure handling, Python code quality questions, AWS/K8s deployment specifics.
Protege product — what you know
  • What it is: LexisNexis flagship AI assistant. Replaced Lexis+ AI Feb 2026. Embedded across Lexis+, Lex Machina, CounselLink+, CourtLink, Intelligize+, PatentSight+.
  • Architecture: Graph RAG (knowledge graph traversal post-vector retrieval — authority ranking). 4-agent framework: orchestrator, legal research, web search, document research + reflection + Shepard's citation verifier.
  • Model stack: Model-agnostic. 14+ models in production. Claude 4, GPT-5, GPT-4o, fine-tuned Mistral (distillation). AWS (Anthropic) + Azure (OpenAI).
  • Scale: 200B+ interconnected legal docs. 4M new docs/day. 2,000+ technologists, ~200 data scientists.
  • CAO thesis: Min Chen — RAG + knowledge graph + distillation + agentic workflows where user can see reasoning.
Anchor stories — trigger → numbers
1 · NEWSRX — self-hosted LLM + multi-agent eval
Sole contributor. vLLM + Qwen3-32B on RunPod. 950+ doc throughput. Built 4-blinded multi-agent eval harness: agents score independently, enrichment agent cross-correlates with HHEM + BERTScore. 27 iterations, 1 variable each. Result: 18% grounding ↑ · 83% human preference · GxP audit trail.
→ Protege bridge: same multi-agent + hallucination mitigation architecture at legal AI scale.
2 · AMGEN — production GenAI in citation-critical domain
First GxP-validated GenAI at Amgen. FDA regulatory submissions — CTD Modules 3, 4, 5. Multi-agent architecture with human signatory accountability. 400+ user stories for AI behavior. Guardrail design + refusal specs. Result: 40% drafting time ↓ · Board expansion after production milestone · Production-grade, not a demo.
→ Protege bridge: legal AI and FDA regulatory AI share the same core problem — citations must be verifiable.
3 · INVISTICS — supervised ML at hospital scale
NIH SBIR $2.1M. Drug diversion ML deployed 300+ hospitals. 96% accuracy. Hybrid feature selection: machine-optimized + domain expert validated. EHR integration across Epic, Cerner, AllScripts, Meditech. Result: 6–8 months faster than manual audits · Maintained accuracy across all EHR diversity.
→ Protege bridge: production ML that ships at scale and stays accurate across data diversity.
Questions to ask them
  • Gaurav: "The knowledge graph traversal post-retrieval for authority ranking — is that happening at retrieval time or as a reranking step? That affects where I'd instrument eval."
  • Gaurav: "For new AI feature rollout — am I primarily extending the existing agent framework or greenfield for a new capability?"
  • Thiyagarajan: "In the multi-agent framework, what are the current failure modes you're most actively solving — loops, context overflow, or citation drift?"
  • Thiyagarajan: "Is the distillation pipeline (teacher→student for smaller models) already in production, or is that the roadmap I'd be building toward?"
  • Both: "Success in 90 days — is there a specific system needing immediate attention, or broader orientation across Protege features?"
2026 LexisNexis AI pressure themes
  • Model agnosticism: routing logic, Best Fit mode, swapping foundation models without pipeline rewrites.
  • Distillation pipeline: teacher→student to reduce cost + latency at scale for high-frequency legal tasks.
  • Personalization: Henchman DMS integration — firm's own matter history as retrieval source. Privacy by design.
  • Agentic maturity: moving from user-observable workflows to 15–20% fully autonomous task completion.
  • Global rollout: Lexis+ with Protege US GA Feb 2026 — international through 2026. Scale + localization pressure.
Metric stack
18% factual grounding ↑ · 27 iterations · 83% human preference · 950+ doc throughput · 400+ user stories · 40% drafting time ↓ · 96% ML accuracy · 300+ hospitals · 4-person plan solo · GxP audit trail
Key differentiators — weave in, don't list
Self-hosted LLM in production
vLLM + Qwen3-32B on RunPod. Not just API wrappers — built the infrastructure, owns the stack.
Multi-agent eval harness built from scratch
4-blinded agents + enrichment agent. Designed from requirements, not assembled from a framework.
Hallucination = liability, not a score
Pharma FDA + SaMD — citation accuracy is regulatory sign-off. Operates at the same stakes as legal AI.
96% accuracy at 300+ hospital scale
Invistics: maintained across Epic/Cerner/AllScripts/Meditech data diversity. Scale and accuracy together.
Sole contributor delivery
NewsRx: 4-person staffing plan, 8 weeks, shipped beyond contracted scope. No team. No playbook.
Production, not demos
Amgen: Board expansion triggered by production milestone. Optinosis SaMD: clinical sign-off on schedule.
Abbreviations — spell-outs
Abbrev.Spell-out / use
RAGRetrieval-Augmented Generation
GxPGood Practices (FDA/ICH validated AI)
SaMDSoftware as a Medical Device
CTDCommon Technical Document (FDA submission)
HHEMHughes Hallucination Evaluation Model
RAGASRetrieval-Augmented Generation Assessment
MCPModel Context Protocol (Anthropic)
CAOChief AI Officer (Min Chen, LexisNexis)
DMSDocument Management System (Henchman integration)
CTHContract to Hire
RELXParent company of LexisNexis (London-listed)
Watch-outs
  • 60 minutes — they will go deep. Budget: 10 min intro, 35 min technical, 10 min your questions. Don't rush anchor stories — Thiyagarajan will drill on details.
  • Thiyagarajan is the bar-raiser. He brings to ALL interviews. Be specific on architecture — don't generalize. He will ask how it works, not just what it does.
  • Legal domain — bridge, don't fake. "Legal AI and pharma regulatory AI share the same core problem — citations must be verifiable." Don't pretend to know case law specifics.
  • Knowledge graph specifics. You know Graph RAG conceptually and authority ranking design pattern. LexisNexis-specific implementation is what you're here to build.
  • Growth mindset signal. Gaurav is 6 months in, building a new team. Show you can figure things out without a playbook — NewsRx and Amgen are both proof points.