Azure — Novant's Cloud Backbone
Azure Data Factory (ADF) — ETL/ELT orchestration. Pulls data from Epic Clarity/Caboodle. Equivalent to Informatica (which Novant is replacing with ADF). Think: the pipeline scheduler and connector hub.
Azure Synapse Analytics — Enterprise data warehouse + analytics. Where structured data lands after ADF pulls it. Enables SQL analytics at scale. Mayur built the multi-terabyte platform here.
Azure Data Lake Storage (ADLS) — Raw and curated data storage. The lake layer. Structured (Synapse) and unstructured data both live here.
Azure DevOps — CI/CD pipelines, sprint boards, version control (git repos). William's team uses this for ML pipeline code promotion (DEV → QA → PROD).
Azure Machine Learning (AML) — Optional ML workflow orchestration. May coexist with Databricks. Provides model registry, compute clusters, automated ML.
Databricks — ML Compute Engine
What it is — Managed Apache Spark platform. Notebooks + jobs + pipelines in one environment. Primary where William's team writes and runs ML code.
PySpark — Python API for Spark. Enables distributed processing of billions of rows (patient labs, vitals, meds, demographics). William used this to process clinical data for the Primary Aldosteronism model.
Databricks Jobs — Scheduled pipelines. William's 30k daily predictions to Epic run as a Databricks job.
Delta Lake — Storage format used with Databricks. ACID transactions on the data lake. Versioned data for auditability — relevant for HIPAA audit trails.
Unity Catalog — Databricks data governance layer. Access controls, lineage, PII tagging. Directly relevant to Bishop's governance background.
MLflow — Model Lifecycle
Experiment Tracking — Logs hyperparameters, metrics, and artifacts for every training run. William built custom MLflow classes using hyperopt for hyperparameter tuning.
Model Registry — Centralized store of versioned models with lifecycle stages: Staging → Production → Archived. The governance control point before a model goes live.
Model Serving — REST endpoint for real-time inference OR batch scoring jobs. William's 30k/day predictions run via batch scoring, not real-time.
MLflow Projects — Packaging standard for reproducible ML code. Enables the CI/CD pipeline William built (pytest, black, flakehell, poetry).
Model Cards — Not native MLflow, but best practice: document performance across demographics, training data, intended use. William assesses demographic bias (sex, race, age) — model cards capture this.
How Bishop's Background Translates to This Stack
| Stack Component | Bishop Analogue | Bridge Statement |
| Azure Data Factory | Informatica + SSIS pipelines (via Novant context); Evernorth Epic ETL work | "At Evernorth I mapped 147 Epic reports to fix the data pipeline before it could serve AI. Same problem ADF solves at the ingest layer." |
| Databricks/PySpark | Cloudera Hadoop (Alliance Mobile Financial); distributed ML pipelines at Wolters Kluwer | "I've operated at Hadoop scale. Databricks is a cleaner abstraction over Spark — the distributed data processing model is familiar." |
| MLflow Model Registry | GxP validation lifecycle at Amgen (IQ/OQ/PQ); model governance at Wolters Kluwer | "The MLflow registry is a lighter version of the GxP validation lifecycle I ran at Amgen. Staging → Production maps to IQ/OQ/PQ. I'd add mandatory model cards and a clinical stakeholder sign-off step before promotion." |
| Azure DevOps CI/CD | JIRA + Agile at AVOXI; SAFe delivery at Amgen; code promotion at J&J MedTech | "I've managed code promotion from DEV→QA→PROD in regulated environments. The Azure DevOps pipeline is the same discipline." |
| Unity Catalog / Data Governance | KPI library + data dictionary at Evernorth; AI governance catalog at J&J | "At Evernorth I built the governed data dictionary and KPI library from the ground up. Unity Catalog is the platform expression of that same work." |
What Each Interviewer Will Probe
Mayur (Platform owner) will ask about:
- How you'd design Epic data extraction in ADF
- Governance controls over data access tiers (who sees what)
- Azure architecture patterns for healthcare HIPAA compliance
- How you handle data quality upstream of ML models
William (ML practitioner) will ask about:
- MLOps rigor — how you manage model versions and approvals
- How you detect and handle model drift in clinical context
- Production monitoring — what signals tell you a model is degrading
- How you validate across demographic subgroups (he does this)
Don't say: "I haven't used MLflow specifically."
Say instead: "My MLOps discipline is framework-agnostic. I've built model governance pipelines across toolchains. MLflow is the standard I'd use here — the model registry and experiment tracking patterns map directly to validation workflows I've run in GxP environments."
Interview Language
"The model registry approval gate is where governance happens — I'd add a mandatory clinical stakeholder review before any model transitions to Production stage."
"Delta Lake's version history gives you the audit trail HIPAA requires — I'd architect the pipeline so every prediction batch is traceable to the exact model version and data snapshot that produced it."
"The Epic → ADF → Data Lake pattern is the data contract that everything else depends on. At Evernorth I learned that if that layer is noisy, the models are wrong — regardless of how well they're tuned."