Multi-agent workflow reliability?
Orchestrator owns state; individual agents stateless + idempotent. Explicit failure contracts — structured result with confidence + failure reason. Reflection agent reviews before delivery (same pattern Protege uses). NewsRx: 4-blinded agents + enrichment agent cross-correlating — extend that here.
Design a verifiable citation system?
Generator and verifier must be separate — same model cannot reliably verify its own output. Citation grounding: every cited case runs through retrieval verification before response finalized. Force structured output (JSON with case_id, jurisdiction, authority_level). Fallback: explicit "could not verify" — never silently omit.
Hallucination in production?
Upstream: verbatim-only constraint for citations/numerics + explicit refusal spec when source unavailable. Architecture: Shepard's pattern — independent citation verification agent. Monitoring: HHEM + BERTScore as pipeline gates, not post-hoc. NewsRx lesson: hallucinations cluster at knowledge gaps — test where data is thin.
Handle agent loops / failures?
Max step budget per agent + explicit termination conditions. Agent returns structured ToolResult with success/failure/reason — orchestrator decides retry vs. escalate vs. fallback. Context window: summarize intermediate results before passing; don't accumulate raw traces.
LLM eval methodology?
Metric set: factual grounding (primary), BERTScore (semantic), HHEM (hallucination), citation accuracy, preference rate. 27-iteration controlled experiments, 1 variable per iteration, full audit trail. RAGAS for RAG-specific (faithfulness, answer relevancy, context recall).
Self-hosted vs. API LLMs?
Self-hosted: data sovereignty, no per-token cost at scale, full control, can fine-tune — GPU infra overhead. API: speed to market, model flexibility (swap GPT-5→Claude 4 without infra change). NewsRx: vLLM + Qwen3-32B on RunPod — eliminated API costs at 950+ doc throughput. LexisNexis hybrid is correct: proprietary legal data on-prem, general AI queries via API.