Cloud AI Stack Comparison — AWS / Azure / GCP / Open Source John Holstein  |  2026-05-18  |  Open Source = Live NewsRx + Optinosis stacks

Capability Layer ☁ AWS ☁ Azure ☁ Google Cloud ⚙ Open Source
(NewsRx + Optinosis — live)
Foundation Models / LLM AccessManaged model endpoints Amazon Bedrock Claude 3.5/3/Haiku, Llama 3.x, Mistral, Titan, Amazon Nova Pro/Lite; unified API; serverless inference; Batch API (50% discount); Guardrails for safety Azure OpenAI Service GPT-4o, GPT-4 Turbo, o1, o3-mini, Embeddings (text-embedding-3); private Azure tenant; same OpenAI API contract; Content Safety filters; DALL-E 3 for multimodal Vertex AI Model Garden Gemini 2.0 Flash/Pro, Claude 3.5 (via Model Garden), Llama 3.x (open), PaLM 2; Gemini API; Model Garden for 100+ models; serverless or dedicated endpoints vLLM + Qwen3-32B LIVE Self-hosted on RunPod H100 SXM 80GB or Modal serverless; PagedAttention for throughput; Anthropic Claude Haiku 3.5 Batch API fallback (50% discount); 950 docs/hr throughput
ML Training PlatformManaged training jobs, compute Amazon SageMaker Training jobs (p4d/p5 instances, H100/A100); SageMaker Studio; HyperPod for distributed training; Autopilot (AutoML); Processing jobs for feature engineering Azure Machine Learning Compute clusters (NC/ND series, H100/A100); AML Studio; AutoML; Responsible AI dashboard; Training pipelines via YAML; MLflow integration native Vertex AI Training Custom training (GPU/TPU v4/v5); pre-built containers (PyTorch, TF, XGBoost); HyperParameter tuning (Vizier); Vertex AI Workbench; Colab Enterprise PyTorch + Axolotl LIVE Optinosis: PyTorch + Hugging Face Transformers for supervised ML on SEER/CMS data; Axolotl (production fine-tuning, DPO+LoRA, multi-GPU); Unsloth (iteration, single GPU); QLoRA for large models on 24GB GPU
Model Serving / InferenceReal-time + batch endpoints SageMaker Endpoints Real-time endpoints; Serverless inference (scale-to-zero); Async inference for long jobs; Multi-model endpoints; SageMaker Batch Transform; JumpStart for prebuilt model deployment Azure ML Online Endpoints Managed online endpoints (real-time); Batch endpoints; Kubernetes endpoints (BYOC); Azure Container Apps for lightweight serving; Model catalog deployments (1-click) Vertex AI Prediction Online prediction endpoints; Batch prediction; Model Armor for governance; Explainable AI (SHAP/IG); pre/post-processing containers; regional endpoint routing vLLM + Modal LIVE Modal serverless functions: scales to zero, per-second billing, ~$5-50/day; RunPod H100 SXM for development; 20,000 doc/day batch capacity; Dagster triggers inference jobs
Vector DB / RAG RetrievalSemantic + hybrid search Amazon OpenSearch Service FAISS-based k-NN vector search; hybrid BM25 + vector; Bedrock Knowledge Bases (managed RAG); Aurora PostgreSQL with pgvector; MemoryDB (Redis vector); Amazon Kendra (enterprise search) Azure AI Search Hybrid retrieval: dense vector + BM25 merged via RRF; Semantic Ranker (cross-encoder reranking); integrated with Azure OpenAI for native RAG; Skillsets for NLP enrichment in indexer; Integrated vectorization Vertex AI Search Managed RAG (Vertex AI RAG Engine GA 2025); AlloyDB with pgvector; Vertex AI Feature Store; Matching Engine (ScaNN for ANN at scale); grounding via Google Search or corpus pgvector + pgvectorscale LIVE NewsRx: pgvector on Render PostgreSQL (zero incremental cost); DiskANN index; Qdrant (mid-term: SIMD Rust, 8ms p50, HNSW filtered ANN); Weaviate considered (BM25+vector RRF); Pinecone rejected (cost model)
Document ExtractionOCR, forms, unstructured → structured Amazon Textract OCR + ML extraction from PDFs, images; Forms (key-value pairs); Tables; Queries API (targeted extraction); Signatures; Expense and Identity docs; Lending AI (specialized); async for large batches Document Intelligence Formerly Form Recognizer; prebuilt models (invoices, receipts, ID, health insurance cards, tax W2); custom models trained on your layouts; confidence scores + bounding boxes; JSON output; batch analysis Document AI Processor library (Invoice, Contract, Expense, Identity, Lending, Healthcare); Enterprise DocAI for custom processors; Workbench for labeling; output: structured Document proto with entity extraction PubMed XML + custom Python LIVE NewsRx: XML fetch from PubMed NCBI API → Python parsing → bronze layer; Optinosis: EHR structured + imaging multimodal. No managed OCR service — purpose-built ingestion for known structured sources
Pipeline OrchestrationDAGs, asset lineage, scheduling AWS Step Functions + MWAA Step Functions: serverless workflow (States Language JSON/YAML); MWAA: managed Apache Airflow 2.x; SageMaker Pipelines (ML-specific DAGs); EventBridge for scheduling; Glue Workflows for ETL Azure Data Factory + AML Pipelines ADF: ETL/ELT pipelines, 90+ connectors, visual designer; Azure ML Pipelines: ML-specific steps in Python SDK; Azure Databricks Workflows; Logic Apps for event-driven; Managed Airflow in ADF Cloud Composer + Vertex Pipelines Cloud Composer: managed Airflow 2.x (GKE-based); Vertex AI Pipelines: Kubeflow Pipelines v2 SDK; native artifact lineage; KFP components; Cloud Scheduler for cron; Dataflow for streaming Dagster + Modal LIVE NewsRx: Dagster (software-defined assets, asset lineage, Components GA Oct 2025); self-hosted on Render worker; Modal handles serverless GPU steps; Prefect considered (alternative); Airflow explicitly rejected
Experiment Tracking / Model RegistryRuns, metrics, artifacts SageMaker Experiments + Model Registry SageMaker Experiments: run tracking, metrics, parameters, artifacts; Model Registry: versioned models, approval workflow, cross-account; SageMaker Model Cards for governance; S3 for artifact storage Azure ML Experiments + Registry Native MLflow integration (log runs, metrics, params); Azure ML Model Registry (versioned, with tags, stage transitions); Responsible AI dashboard per model; environment and dataset versioning Vertex AI Experiments + Model Registry Vertex AI Experiments: TensorBoard-backed; Metadata API for lineage; Vertex AI Model Registry: versioned, with labels and deployment history; Vertex AI Evaluation for systematic model comparison MLflow LIVE Both NewsRx and Optinosis: MLflow for run tracking, metric logging, artifact storage, model registry; runcard documentation system for IEC 62304 / 21 CFR Part 11 traceability; 27 documented prompt iterations at NewsRx
Monitoring / Model ObservabilityDrift, performance, data quality SageMaker Model Monitor + CloudWatch Model Monitor: data quality, model quality, bias drift, feature attribution drift; CloudWatch dashboards + alarms; SageMaker Clarify for explainability and bias; real-time or scheduled monitoring jobs Azure ML Model Monitoring + Azure Monitor AML monitoring: data drift, prediction drift, feature distribution tracking; Azure Monitor for infra metrics + alerts; Application Insights for app telemetry; Responsible AI dashboard for fairness/explainability Vertex AI Model Monitoring + Cloud Monitoring Feature skew and drift detection; prediction skew vs. training data; email/PagerDuty alerts on threshold breach; Cloud Monitoring for infra; Vertex AI Evaluation for systematic evals HHEM + BERTScore + PSI LIVE NewsRx: HHEM-2.1-Open (hallucination scoring, target ≥0.90), BERTScore F1 (target ≥0.90), Flesch-Kincaid via textstat; custom PSI monitoring for distributional shift; Dagster observability for pipeline health; Streamlit SME review interface
Data Catalog / GovernanceLineage, discovery, compliance AWS Glue Data Catalog + Lake Formation Glue Data Catalog: Hive-compatible metadata store, 1M+ tables; Lake Formation: column/row-level security, data sharing, governed tables; Macie for PII detection in S3; DataZone for data marketplace Microsoft Purview + Unity Catalog Purview: data map, lineage, classification, sensitivity labels, policy; Databricks Unity Catalog (Azure Databricks): column-level security, row filters, audit log; Azure Data Lake RBAC via POSIX ACLs Dataplex + BigQuery Data Catalog Dataplex: unified data governance across lakes, warehouses, marts; auto data discovery + lineage; data quality rules as code; BigQuery Data Catalog: tag templates, policy tags, column-level security Alation + GxP/ALCOA+ Framework LIVE J&J MedTech (Bishop): Alation semantic layer, data dictionary, lineage tracking, FDA+EU MDR dual compliance; Optinosis: MLflow lineage, IEC 62304, 21 CFR Part 11, HIPAA, ALCOA+ principles; C2PA + PROV-O at NewsRx; 301 governed artifacts
Storage / Data LayerObject, warehouse, OLTP S3 + Redshift + RDS S3: object storage (data lake foundation); Redshift: columnar OLAP; RDS PostgreSQL/Aurora; Glue for ETL; S3 + Delta Lake / Iceberg for lakehouse; DynamoDB for NoSQL; ElastiCache for caching Azure Blob + Azure SQL + Synapse Blob Storage: ADLS Gen2 for data lake; Azure SQL Database; Synapse Analytics: unified analytics (Spark + SQL pools + Pipelines); Cosmos DB (NoSQL); Cache for Redis; Delta Lake native in Synapse GCS + BigQuery + Cloud SQL GCS: object storage; BigQuery: serverless columnar analytics (pay per query); Cloud SQL (PostgreSQL/MySQL); Spanner (globally distributed OLTP); Bigtable (NoSQL at scale); Memorystore for caching PostgreSQL + Medallion Architecture LIVE NewsRx: PostgreSQL on Render; Bronze/Silver/Gold medallion layers; idempotent append-only ingestion; feed delivery audit logging; train/eval split enforced at pipeline level; pgvector on same instance (zero added cost)
Serverless / Scale-to-Zero ComputeEvent-driven, no idle cost AWS Lambda + Fargate + Batch Lambda: 15-min max, 10GB memory, 6 vCPU; Fargate: containerized, no server management; Batch: managed job queues for HPC/ML; EventBridge triggers; Bedrock serverless inference eliminates GPU management Azure Functions + Container Apps Functions: event-driven, consumption plan (scale-to-zero); Container Apps: serverless containers with KEDA autoscaling; Azure ML Serverless Endpoints (scale-to-zero per request); Logic Apps for workflow automation Cloud Functions + Cloud Run Functions gen2: 60-min timeout, 16GB RAM; Cloud Run: containerized, concurrency-based autoscaling, scale-to-zero; Vertex AI Serverless Prediction; Cloud Run GPUs (NVIDIA L4/T4, 2024) Modal LIVE NewsRx: Modal serverless GPU functions; per-second billing; scales to zero between runs; ~$5-50/day depending on model; $30/month free tier covers dev; native Python; H100/A100/T4 available; async batch support
Fine-Tuning InfrastructurePEFT, LoRA, DPO, adapters SageMaker HyperPod + JumpStart HyperPod: resilient distributed training clusters; JumpStart: 1-click fine-tune (LoRA, QLoRA) on Llama, Falcon, etc.; SageMaker Training: bring own PEFT script; S3 for dataset + checkpoints Azure ML Fine-Tuning + AI Studio Azure AI Studio: fine-tune Llama, Phi-3, Mistral with LoRA; managed compute; AML custom training with Axolotl/Unsloth in custom containers; RAFT (fine-tune on synthetic RAG data) supported Vertex AI Supervised Tuning + Model Garden Gemini fine-tuning (supervised + RLHF); open model fine-tuning via Model Garden (Llama); Vertex AI Training custom jobs (bring Axolotl container); TPU pods for large-scale fine-tuning Axolotl + Unsloth + LoRA/DPO LIVE Optinosis: Axolotl (production; multi-GPU; DPO native; YAML config); Unsloth (iteration + validation; faster, single GPU); LoRA (10-100MB adapters); QLoRA (large models on 24GB GPU); DPO (40% compute savings vs. RLHF)
UI / Application LayerFront-end for AI apps Amplify + Bedrock Agents AWS Amplify: full-stack React/Next.js hosting; Bedrock Agents: conversational AI with action groups; Bedrock Knowledge Bases for RAG-backed chat; AppSync (GraphQL API); CloudFront CDN Azure Static Web Apps + Bot Service Static Web Apps: React/Next.js with serverless APIs; Azure Bot Service: multi-channel conversational AI; Azure AI Studio Playground; Power Apps for low-code; Teams integration for enterprise chat Firebase + Vertex AI Agent Builder Firebase Hosting + Cloud Functions; Vertex AI Agent Builder (formerly Dialogflow CX): conversational agents; Gemini in Workspace; Looker Studio for BI dashboards; Cloud Run for custom web apps Streamlit + Render LIVE Both NewsRx and Optinosis: Streamlit for SME review UI, analytics dashboards, investor/client demos; hosted on Render; Pydantic v2 for schema validation; MCP (Model Context Protocol) for pipeline automation at Optinosis

Cloud AI Stack — Differentiators, Bridge Language & Open Source Advantage John Holstein  |  2026-05-18

AWS — When to Lead With It

  • Bedrock is the strongest managed LLM inference + RAG story (Knowledge Bases, Guardrails)
  • SageMaker is the most mature ML platform; battle-tested at enterprise scale
  • AWS wins on breadth of data services (Redshift, Glue, Lake Formation, Macie)
  • Your credentials: J&J MedTech (S3, Redshift, Databricks on AWS); Amgen (AWS Redshift, IQVIA); Invistics (Oracle → AWS S3)
  • Differentiator: Lambda + Step Functions for event-driven ML pipelines; no Kubernetes required
  • Watch-out: Vendor lock-in on Step Functions States Language; Bedrock model selection narrower than Azure Model Catalog

Azure — When to Lead With It

  • Azure OpenAI is the only way to get GPT-4o/o1 in a private tenant with enterprise SLA — critical for regulated industries
  • Azure AI Search hybrid retrieval (vector + BM25 + semantic ranker) is the most production-ready managed RAG retrieval layer
  • Document Intelligence strongest prebuilt model library for insurance/healthcare forms
  • Azure wins when the client has Microsoft 365 / Teams / Entra ID — seamless enterprise identity
  • Differentiator: Microsoft Purview for unified data governance; Azure ML native MLflow = zero migration friction
  • Watch-out: Azure OpenAI quota allocation can bottleneck; limited non-OpenAI model choice (vs. Bedrock/Vertex)

GCP — When to Lead With It

  • Vertex AI is the tightest integration between training, serving, evaluation, and experiment tracking
  • BigQuery serverless analytics is unmatched for ad-hoc large-scale ML feature computation
  • Gemini 2.0 native multimodal (text, image, video, audio, code) with long context (1M tokens)
  • GCP wins on TPU access for large-scale fine-tuning; Google Search grounding for RAG
  • Differentiator: Vertex AI RAG Engine (GA 2025) tightest managed RAG; Kubeflow Pipelines v2 native in Vertex — your Optinosis Kubeflow experience maps directly
  • Watch-out: Smaller ecosystem of third-party tools vs. AWS; less insurance/healthcare enterprise penetration

Open Source — Your Competitive Edge

  • You're building this live — not classroom exercises. NewsRx and Optinosis are production systems
  • Self-hosted LLM (Qwen3-32B / vLLM) eliminates API cost; at 20K docs/day, saves $400-840/month vs. Claude API
  • Dagster asset lineage = richer governance than most managed orchestrators
  • MLflow across both projects = portable experiment tracking independent of cloud
  • Differentiator: You know what cloud services abstract away — you've built the primitives. This is rare.
  • Interview power move: "I've implemented this from primitives; I can operate any cloud platform because I understand what each service is replacing."

Your Open Source Stack → Cloud Service Bridge (Use in Any Interview)

What You Built (OSS) AWS Equivalent Azure Equivalent GCP Equivalent
vLLM + Qwen3-32B (self-hosted inference) Bedrock (Claude, Llama) or SageMaker JumpStart endpoint Azure OpenAI endpoint or AML managed endpoint Vertex AI endpoint (Gemini or Model Garden)
Dagster (software-defined assets, lineage) AWS Step Functions + SageMaker Pipelines Azure ML Pipelines + ADF Vertex AI Pipelines (KFP v2) — Kubeflow native (Optinosis)
Modal (serverless GPU, scale-to-zero) SageMaker Serverless Inference + Lambda for orchestration Azure ML Serverless Endpoints + Container Apps Cloud Run (GPU L4/T4) + Vertex AI Serverless Prediction
pgvector + pgvectorscale (vector search) Aurora pgvector or OpenSearch k-NN; Bedrock Knowledge Bases Azure AI Search (hybrid retrieval + semantic ranker) AlloyDB pgvector or Matching Engine (ScaNN); Vertex RAG Engine
Qdrant (HNSW filtered ANN, 8ms p50) OpenSearch FAISS k-NN with filtered queries Azure AI Search with vector filter expressions Vertex Matching Engine with restrict tokens for filtering
MLflow (experiments, model registry) SageMaker Experiments + Model Registry Azure ML native MLflow (identical API, zero migration) Vertex AI Experiments + Model Registry
HHEM + BERTScore (hallucination, semantic eval) SageMaker Clarify (bias/explainability); no native hallucination scoring Azure ML Responsible AI dashboard (groundedness metric) Vertex AI Evaluation (pointwise + pairwise eval; groundedness)
Axolotl + LoRA/DPO (fine-tuning) SageMaker HyperPod + JumpStart fine-tuning (LoRA prebuilt) Azure AI Studio fine-tuning (LoRA; Llama, Phi-3, Mistral) Vertex AI Supervised Tuning; custom training job with Axolotl container
PostgreSQL medallion (Bronze/Silver/Gold) S3 + Glue (Bronze) → Redshift (Gold); Delta Lake on S3 ADLS Gen2 + Synapse Analytics; Delta Lake on Synapse GCS (Bronze) → BigQuery (Gold); Delta Lake on Dataproc
Streamlit (SME review UI, dashboards) SageMaker Ground Truth (labeling); Amplify for custom UI Azure Static Web Apps + Power Apps (low-code) Firebase Hosting + Looker Studio; Vertex AI Agent Builder playground
Alation (semantic layer, data catalog) AWS Glue Data Catalog + DataZone Microsoft Purview (unified governance + lineage) Dataplex + BigQuery Data Catalog
Kubeflow + Kubernetes (Optinosis) SageMaker HyperPod (Kubernetes-based) or EKS + Kubeflow AKS + Kubeflow or Azure ML Kubernetes compute Vertex AI Pipelines (KFP v2 = Kubeflow native) — direct map
RunPod H100 SXM (GPU dev/inference) EC2 p5.48xlarge (H100) or SageMaker ml.p5 instances Azure NDH100v5 series (H100 SXM) Vertex AI a3-highgpu (H100) or custom training with H100 nodes
GxP / ALCOA+ / IEC 62304 governance SageMaker Model Cards + CloudTrail audit; no native GxP Azure ML + Purview; no native GxP (you bring the framework) Vertex AI + Dataplex; no native GxP (you bring the framework)

Interview Phrases — Bridging Your Stack to Any Cloud

"I've built the primitives that these managed services abstract — self-hosted vLLM, custom medallion pipelines, OSS vector search. That means I can operate any cloud platform because I understand what each service is replacing."
"My Kubeflow experience at Optinosis maps directly to Vertex AI Pipelines — KFP v2 is Kubeflow native. If this is a GCP shop, I'm working in familiar territory at the pipeline level."
"Azure AI Search is the managed version of what I've built with pgvector + Qdrant — hybrid retrieval, filtered ANN, reranking. The architecture is the same; the operational overhead is different."
"My MLflow experience is cloud-portable — same API whether I'm on AWS, Azure ML (which natively wraps MLflow), or self-hosted. I won't need a migration ramp."
"SageMaker Clarify does the explainability layer; I've built equivalent evaluation using HHEM and BERTScore in production. I can speak to what each cloud service covers and what it doesn't."
"For regulated AI, none of these cloud platforms provide GxP out of the box — you bring that framework. I've delivered GxP-validated AI at Amgen under ICH guidelines and built ALCOA+-aligned governance at Optinosis from the ground up."

Quick Reference — Which Stack to Name When

If the client says...Lead with...
"We're an Azure shop / Microsoft stack"Azure AI Search for RAG, Azure OpenAI for LLM, Document Intelligence for extraction; mention Purview for governance; your OSS stack maps to all three
"We use AWS / Bedrock"SageMaker for training, Bedrock Knowledge Bases for RAG; your S3/Redshift/Databricks experience at J&J + Amgen; Invistics AWS S3 migration
"We're on GCP / BigQuery"Vertex AI Pipelines (your Kubeflow/Kubernetes = direct map); BigQuery as Gold layer (your PostgreSQL medallion maps); Vertex RAG Engine
"We're cloud-agnostic / multi-cloud"Your OSS stack is cloud-portable by design: MLflow, Dagster, pgvector all run anywhere; Kubernetes at Optinosis = deploy to EKS, AKS, or GKE
"We need to control costs" / "Azure spending is high"Self-hosted LLM story: Qwen3-32B on Modal saves $400-840/month vs. Claude API at NewsRx scale; Modal serverless eliminates idle GPU cost
"We need GxP / FDA-validated AI"Amgen: GxP-validated GenAI under ICH guidelines; Optinosis: IEC 62304 + 21 CFR Part 11 + ISO 14971; this is framework you bring to any cloud
"We use Snowflake"Optum Medicaid TPA analytics: unified Snowflake source across multiple business lines — direct production experience
"What's your fine-tuning experience?"Axolotl + LoRA + DPO at Optinosis; QLoRA for large models on 24GB GPU; Unsloth for iteration; maps to SageMaker JumpStart, Azure AI Studio fine-tuning, Vertex supervised tuning