Arbitration Forums — AI Product Architect — Round 1 Screening Prep

John Holstein  |  KFORCE (Roderick Bethune: 813-552-3845)  |  2026-05-18

Company Intelligence — Arbitration Forums, Inc.

What they areNation's largest P&C insurance inter-company arbitration and subrogation nonprofit. Founded 1943, Tampa FL. Membership-driven: carriers pay per filing, not annual fees.
Scale5,400+ member carriers; 1.1M+ disputes/year; $26.4B in claims processed. Platform: E-Subro Hub® (electronic routing of subrogation demands + arbitration filings).
Tech momentOct 2025: Board approved capital allocation to tech/AI strategy. Dec 2024: Guidewire Cloud integration live. Snowflake confirmed in hiring signals. Active build → adoption phase. 43+ open tech roles.
CITOJohn Shedd (10+ yrs). CISSP, M.S. Cybercrime, B.S. Info Systems. Security-first lens. Prior: Catalina Marketing, HSN. Values: Agile, team quality, business-tech alignment.
AI opportunityDispute outcome prediction, document classification/routing, subrogation demand extraction (RAG), arbitrator assignment optimization, anomaly detection in claim patterns.
StackConfirmed: Snowflake, Guidewire Cloud. JD: Azure AI Search, Azure OpenAI, Document Intelligence, SQL Server. Not confirmed public: Azure vs. AWS preference.
HCL roleSystems integration partner; placing external data science/AI talent to build AF's capability. KFORCE also recruiting (AI Product Architect title). Tampa-based = local relationship.

Critical Framing for Written Responses

  • "Please do not use AI tools when answering" — KFORCE explicitly required authentic responses. These answers must be in your own voice. Use this prep to anchor your stories, then write your own.
  • One-and-done interview: Strong written answers = they proceed directly to interview. Weak or generic answers = out. Every answer must be specific, first-person, tool-named.
  • Lead with hands-on: You built/coded/deployed — not "I architected the strategy for." Name the tool, name the data, name the outcome metric. (Home Depot lesson.)
  • Title bridge: JD says "Data Scientist"; KFORCE says "AI Product Architect." You are both — architect who builds. Frame as: "I design and deliver."

Positioning Statement (Open with This Framing)

"My relevant work combines hands-on ML system delivery in high-data-volume, regulated environments with end-to-end production deployment at scale. At Invistics I built and deployed a supervised ML drug diversion detection system to 300+ hospitals — that engagement had the same data profile as insurance claims arbitration: high-volume structured transactions, unstructured document data, anomaly detection objectives, and HIPAA-grade privacy requirements. That's the direct analog to what AF is building."

Why this lands: It names a real production deployment, draws the explicit analogy to their domain, leads with scale, and establishes healthcare → insurance credibility transfer before Q1 is even answered.


Q1
"Walk me through an AI capability you personally built end-to-end that was integrated into a business process. What did you actually design and implement?"
Primary anchor: Invistics drug diversion detection (Invistics/Wolters Kluwer, 2018–2021)
Most direct analog to AF's problem set. Scale (300+ hospitals) + production ML + HIPAA + workflow integration.
The business problem: Hospitals lose $700M+/year to opioid diversion; manual audits catch it 6–8 months after the fact. We built an ML system to detect it in near real-time using controlled substance transaction data.
What I designed: Supervised ML feature set using hybrid selection — ML optimization cross-validated against a 12-person expert panel (investigators, nurses, physicians, pharmacists, regulators, law enforcement). I assembled and ran the panel, mapped the controlled substance workflow end-to-end, and defined failure modes for feature engineering targets. Risk output: three-tier categorization (high-risk / behavioral pattern / sloppy behavior) mapped to investigator workflow queues.
What I implemented: Integration platform normalizing controlled substance transaction data across four EHR schemas (Epic, Cerner, AllScripts, Meditech) and two ADC systems (BD Pyxis, Omnicell) — Python, SQL, Oracle, ETL/batch processing. Model trained on combined multi-EHR dataset. Output fed directly into investigator queue dashboard, replacing paper audit trails.
Business integration outcome: 96% detection accuracy; diversions identified 6–8 months faster than manual audits; system confirmed existing suspicions AND surfaced unknown diverters who proved out under investigation; deployed to 300+ US hospitals. HIPAA + GDPR compliant throughout.
Tools named: Python, supervised ML, feature engineering, Oracle → AWS S3, ETL, SQL, batch processing, NIH-funded ($2.1M SBIR grant delivery)
Q2
"How did you handle data ingestion, data quality, and governance for that solution?"
Primary anchor: Invistics (continuing Q1 story) + J&J MedTech for governance depth
Start with Invistics ingestion/quality, then reference J&J for governance architecture.
Ingestion challenge: Each hospital's EHR vendor had different controlled substance transaction schemas, timestamps, user identifier formats, and ADC integration patterns. I built an abstraction layer in Python that normalized these into a common controlled substance event schema regardless of source system. No per-facility customization required — new hospitals onboarded by data mapping only.
Data quality: Three categories — completeness (missing dose records common in Meditech environments; handled with configurable thresholds and flagging), accuracy (cross-validated dispense-to-waste reconciliation records against independent ADC logs), and consistency (timestamp normalization across 24hr vs. 12hr formats, facility timezone handling). Quality gates ran before records entered the ML pipeline.
Governance (J&J MedTech reference for depth): At J&J I built the full governance stack — Alation semantic layer for data cataloging, data dictionary, lineage tracking from SAP/JDE source feeds through AWS S3 → Databricks transformation → Redshift gold layer. GxP compliance required full audit trail on every schema change, zero production defects over 2 years under FDA and EU MDR dual compliance. That governance architecture is what I bring to the Invistics deployment model as well — every pipeline change documented and version-controlled.
HIPAA/PHI specifically: PII isolation at ingestion — patient identifiers were hashed at the source; ML models ran on de-identified transaction features; investigator workflow surfaced identifiers only after alert threshold crossed and HIPAA-compliant access protocol triggered. No raw PHI traversed to central server.

Arbitration Forums — AI Product Architect — Round 1 (Back)

John Holstein  |  2026-05-18
Q3
"How did you make the solution production-ready (monitoring, security, scalability)?"
Primary anchor: Invistics (300+ hospitals = real production scale) + Amgen (regulatory-grade monitoring)
Security: HIPAA Business Associate Agreements with each hospital; no raw PHI transmitted to central infrastructure — ML model operated on de-identified feature vectors; facility-specific data never co-mingled. GDPR compliance layer for international partnerships. Access control: role-based, investigator credentials required to view alert details.
Monitoring: Two-track monitoring system. (1) Performance monitoring: investigator feedback loop — every alert opened was tracked to outcome (confirmed diversion vs. false positive); false positive rate and confirmed rate tracked quarterly; model performance dashboard built for operations team. (2) Input monitoring: data feed health checks — if a hospital's EHR data was missing, delayed, or schema-changed, automated alerting prevented silent model degradation.
Scalability: Designed for horizontal scale — normalization layer abstracted EHR schema differences, so adding a new hospital was a configuration and data mapping exercise, not a code change. Batch processing architecture handled hospital sizes from 200-bed community to 1,000-bed academic medical center. Oracle → AWS S3 migration enabled elastic scaling of data infrastructure as the deployment grew.
Amgen addendum (for regulated production rigor): At Amgen, production-readiness meant GxP validation — Installation Qualification, Operational Qualification, Performance Qualification documentation; change control for every code modification; human signatory accountability preserved in the GenAI output chain; IQ/OQ/PQ on the full system before FDA-facing use. I know what it means to build AI that can't fail in production.
Q4
"How have you handled model drift or changes in performance over time?"
Primary anchor: Invistics 3-year production operation + NewsRx controlled experimentation discipline
What caused drift at Invistics: Three sources. (1) EHR version drift — Epic and Cerner release updates that change how they log controlled substance transactions; adapter layer abstracted schema changes so drift was caught at ingestion before reaching the model. (2) Distributional shift — DEA rescheduling events (new drugs added/removed from schedules) changed the underlying data distribution; monitored feature distributions monthly using Python; Population Stability Index (PSI) thresholds triggered review. (3) Concept drift — diversion tactics evolved (novel medications, new diversion methods); investigator label quality degraded as what constituted "suspicious" shifted; addressed through periodic annotation sessions with subject matter experts to relabel edge cases.
Response protocol: Two-track trigger system: performance track (confirmed diversion rate dropped >10% from baseline) OR distribution track (PSI > 0.2 on key features). On trigger: root cause analysis (ingestion change vs. true distributional shift vs. concept drift), targeted feature updates, validation on held-out test set from recent time period, staged rollout with A/B comparison against previous model version.
NewsRx for LLM drift (more current): At NewsRx I established controlled experimentation with 27 documented prompt iterations tracking HHEM hallucination scores and BERTScore across iterations — same discipline: establish a baseline, define a trigger threshold (HHEM drop > threshold = intervention), document every change with before/after metrics. When model versions update (vLLM or Qwen3 releases), re-run the full evaluation suite before any production traffic moves. Same principle as classical model drift monitoring, applied to the GenAI context.
In AF's context: Insurance claim patterns shift with litigation trends, state regulatory changes, and case law. I'd build population monitoring on the key feature distributions (claim amounts, dispute type frequencies, carrier pair patterns) with automated alerts, and establish a semi-annual review cadence with AF's arbitration domain experts — the same expert-grounded revalidation I used at Invistics.

Say This, Not That

AvoidSay Instead
"I architected a solution for...""I built and deployed..." / "I implemented..."
"We designed a governance framework""I built the Alation semantic layer, data dictionary, and lineage tracking"
"I oversaw a team that built...""I led the feature engineering and coded the integration adapters"
"I have experience with Azure AI""I've deployed RAG pipelines in production; I'm mapping those patterns to Azure AI Search — it's architecturally the same retrieval mechanism"
"I'm familiar with the insurance domain""My Invistics deployment had the same data profile as AF's problem: high-volume structured transactions, unstructured docs, anomaly detection, HIPAA compliance"
"I'm very hands-on"[Name a specific tool + what you did with it + the outcome]

Watch-Outs

AI Tool Prohibition (KFORCE)

Roderick explicitly said "please do not use AI tools when answering." Your authentic responses are what triggers the one-and-done interview. Write in your own voice. Use this cheat sheet to anchor stories, then write yourself.

Azure Exposure Gap

Your documented projects are AWS-heavy (AWS S3, Redshift, SageMaker). Don't overclaim Azure hands-on. Frame as: "I've implemented the patterns this JD describes on AWS; the Azure equivalents are architectural translations, not new concepts." Then name the specific Azure tool and what it maps to.

FHIR in a P&C Context

AF is P&C insurance, not a health system. FHIR is an EHR interoperability standard — it's in the JD because the HCL template is generic. Don't lead with FHIR. HIPAA/PHI is the right frame (bodily injury claims contain medical records). Mention FHIR only if directly asked.

Overqualification Signal

COO, 20+ years, VP history. If it comes up: "I'm at this engagement level by choice — I build the systems, I don't just design them. My most impactful work has been as the person who both architects and delivers."

Questions to Ask AF (When You Get the Interview)

  • Snowflake state: "What's the current state of your Snowflake deployment — are you building the data layer or scaling an existing one?"
  • E-Subro Hub as data source: "Is E-Subro Hub the primary training data source, or are there additional claims management systems feeding the AI use cases?"
  • Build vs. adopt: "Have there been any AI POCs at AF that didn't make it to production? What limited them?"
  • Success definition: "What does success look like at 90 days for this role — is it a deployed model, a data architecture decision, or something else?"
  • Team structure: "Who is my primary technical counterpart — data engineering, the ops team, or directly with John Shedd's organization?"
  • Arbitrator panel dynamics: "How are arbitration domain experts currently involved in model validation? That expert-grounded approach was key to my Invistics work."

Key Metrics to Drop Naturally

ProjectMetric to Use
Invistics96% detection accuracy; 6–8 months faster; 300+ hospitals; $2.1M NIH SBIR
Amgen40% CMC drafting time reduction; 400+ user stories to production; Board expansion triggered
NewsRx950+ docs; 18% factual grounding improvement; 88% human preference rate; 27 iterations
J&J MedTech$750K annual cost reduction; zero production defects 2 years; FDA + EU MDR dual compliance
Optum/SnowflakeUnified Snowflake analytics source; Medicaid + TPA claims; 4 team members absorbed to permanent
Cardinal Health$21M SAP cGMP implementation; 80+ radiopharmacies; $300M+ in contract violation recovery

Final Reminder

Every Q answer ends with: production deployment + outcome metric + tool named. The client will proceed to a one-and-done interview if these answers are specific, authentic, and hands-on. Generic = rejected before reading. Specific = interview scheduled.