Azure + Databricks + MLflow — Novant Health Tech Stack  |  Crash Course  |  2026-05-21

Mental Model: Novant's AI stack is three layers: Azure is the cloud infrastructure and governance layer; Databricks is the compute and ML engineering layer; MLflow is the experiment and model lifecycle layer. Mayur (Director AI Platform) owns the Azure infrastructure including Epic data pipelines via Azure Data Factory. William (Data Science Manager) works in Databricks + MLflow daily. The AI Technical Product Architect sits between them — translating product requirements into architecture decisions that both sides can execute on.

Azure — Novant's Cloud Backbone

Azure Data Factory (ADF) — ETL/ELT orchestration. Pulls data from Epic Clarity/Caboodle. Equivalent to Informatica (which Novant is replacing with ADF). Think: the pipeline scheduler and connector hub.
Azure Synapse Analytics — Enterprise data warehouse + analytics. Where structured data lands after ADF pulls it. Enables SQL analytics at scale. Mayur built the multi-terabyte platform here.
Azure Data Lake Storage (ADLS) — Raw and curated data storage. The lake layer. Structured (Synapse) and unstructured data both live here.
Azure DevOps — CI/CD pipelines, sprint boards, version control (git repos). William's team uses this for ML pipeline code promotion (DEV → QA → PROD).
Azure Machine Learning (AML) — Optional ML workflow orchestration. May coexist with Databricks. Provides model registry, compute clusters, automated ML.

Databricks — ML Compute Engine

What it is — Managed Apache Spark platform. Notebooks + jobs + pipelines in one environment. Primary where William's team writes and runs ML code.
PySpark — Python API for Spark. Enables distributed processing of billions of rows (patient labs, vitals, meds, demographics). William used this to process clinical data for the Primary Aldosteronism model.
Databricks Jobs — Scheduled pipelines. William's 30k daily predictions to Epic run as a Databricks job.
Delta Lake — Storage format used with Databricks. ACID transactions on the data lake. Versioned data for auditability — relevant for HIPAA audit trails.
Unity Catalog — Databricks data governance layer. Access controls, lineage, PII tagging. Directly relevant to Bishop's governance background.

MLflow — Model Lifecycle

Experiment Tracking — Logs hyperparameters, metrics, and artifacts for every training run. William built custom MLflow classes using hyperopt for hyperparameter tuning.
Model Registry — Centralized store of versioned models with lifecycle stages: Staging → Production → Archived. The governance control point before a model goes live.
Model Serving — REST endpoint for real-time inference OR batch scoring jobs. William's 30k/day predictions run via batch scoring, not real-time.
MLflow Projects — Packaging standard for reproducible ML code. Enables the CI/CD pipeline William built (pytest, black, flakehell, poetry).
Model Cards — Not native MLflow, but best practice: document performance across demographics, training data, intended use. William assesses demographic bias (sex, race, age) — model cards capture this.

How Bishop's Background Translates to This Stack

Stack ComponentBishop AnalogueBridge Statement
Azure Data FactoryInformatica + SSIS pipelines (via Novant context); Evernorth Epic ETL work"At Evernorth I mapped 147 Epic reports to fix the data pipeline before it could serve AI. Same problem ADF solves at the ingest layer."
Databricks/PySparkCloudera Hadoop (Alliance Mobile Financial); distributed ML pipelines at Wolters Kluwer"I've operated at Hadoop scale. Databricks is a cleaner abstraction over Spark — the distributed data processing model is familiar."
MLflow Model RegistryGxP validation lifecycle at Amgen (IQ/OQ/PQ); model governance at Wolters Kluwer"The MLflow registry is a lighter version of the GxP validation lifecycle I ran at Amgen. Staging → Production maps to IQ/OQ/PQ. I'd add mandatory model cards and a clinical stakeholder sign-off step before promotion."
Azure DevOps CI/CDJIRA + Agile at AVOXI; SAFe delivery at Amgen; code promotion at J&J MedTech"I've managed code promotion from DEV→QA→PROD in regulated environments. The Azure DevOps pipeline is the same discipline."
Unity Catalog / Data GovernanceKPI library + data dictionary at Evernorth; AI governance catalog at J&J"At Evernorth I built the governed data dictionary and KPI library from the ground up. Unity Catalog is the platform expression of that same work."

What Each Interviewer Will Probe

Mayur (Platform owner) will ask about:
  • How you'd design Epic data extraction in ADF
  • Governance controls over data access tiers (who sees what)
  • Azure architecture patterns for healthcare HIPAA compliance
  • How you handle data quality upstream of ML models
William (ML practitioner) will ask about:
  • MLOps rigor — how you manage model versions and approvals
  • How you detect and handle model drift in clinical context
  • Production monitoring — what signals tell you a model is degrading
  • How you validate across demographic subgroups (he does this)
Don't say: "I haven't used MLflow specifically."
Say instead: "My MLOps discipline is framework-agnostic. I've built model governance pipelines across toolchains. MLflow is the standard I'd use here — the model registry and experiment tracking patterns map directly to validation workflows I've run in GxP environments."

Interview Language

"The model registry approval gate is where governance happens — I'd add a mandatory clinical stakeholder review before any model transitions to Production stage."
"Delta Lake's version history gives you the audit trail HIPAA requires — I'd architect the pipeline so every prediction batch is traceable to the exact model version and data snapshot that produced it."
"The Epic → ADF → Data Lake pattern is the data contract that everything else depends on. At Evernorth I learned that if that layer is noisy, the models are wrong — regardless of how well they're tuned."