| Foundation Models / LLM AccessManaged model endpoints |
Amazon Bedrock
Claude 3.5/3/Haiku, Llama 3.x, Mistral, Titan, Amazon Nova Pro/Lite; unified API; serverless inference; Batch API (50% discount); Guardrails for safety
|
Azure OpenAI Service
GPT-4o, GPT-4 Turbo, o1, o3-mini, Embeddings (text-embedding-3); private Azure tenant; same OpenAI API contract; Content Safety filters; DALL-E 3 for multimodal
|
Vertex AI Model Garden
Gemini 2.0 Flash/Pro, Claude 3.5 (via Model Garden), Llama 3.x (open), PaLM 2; Gemini API; Model Garden for 100+ models; serverless or dedicated endpoints
|
vLLM + Qwen3-32B LIVE
Self-hosted on RunPod H100 SXM 80GB or Modal serverless; PagedAttention for throughput; Anthropic Claude Haiku 3.5 Batch API fallback (50% discount); 950 docs/hr throughput
|
| ML Training PlatformManaged training jobs, compute |
Amazon SageMaker
Training jobs (p4d/p5 instances, H100/A100); SageMaker Studio; HyperPod for distributed training; Autopilot (AutoML); Processing jobs for feature engineering
|
Azure Machine Learning
Compute clusters (NC/ND series, H100/A100); AML Studio; AutoML; Responsible AI dashboard; Training pipelines via YAML; MLflow integration native
|
Vertex AI Training
Custom training (GPU/TPU v4/v5); pre-built containers (PyTorch, TF, XGBoost); HyperParameter tuning (Vizier); Vertex AI Workbench; Colab Enterprise
|
PyTorch + Axolotl LIVE
Optinosis: PyTorch + Hugging Face Transformers for supervised ML on SEER/CMS data; Axolotl (production fine-tuning, DPO+LoRA, multi-GPU); Unsloth (iteration, single GPU); QLoRA for large models on 24GB GPU
|
| Model Serving / InferenceReal-time + batch endpoints |
SageMaker Endpoints
Real-time endpoints; Serverless inference (scale-to-zero); Async inference for long jobs; Multi-model endpoints; SageMaker Batch Transform; JumpStart for prebuilt model deployment
|
Azure ML Online Endpoints
Managed online endpoints (real-time); Batch endpoints; Kubernetes endpoints (BYOC); Azure Container Apps for lightweight serving; Model catalog deployments (1-click)
|
Vertex AI Prediction
Online prediction endpoints; Batch prediction; Model Armor for governance; Explainable AI (SHAP/IG); pre/post-processing containers; regional endpoint routing
|
vLLM + Modal LIVE
Modal serverless functions: scales to zero, per-second billing, ~$5-50/day; RunPod H100 SXM for development; 20,000 doc/day batch capacity; Dagster triggers inference jobs
|
| Vector DB / RAG RetrievalSemantic + hybrid search |
Amazon OpenSearch Service
FAISS-based k-NN vector search; hybrid BM25 + vector; Bedrock Knowledge Bases (managed RAG); Aurora PostgreSQL with pgvector; MemoryDB (Redis vector); Amazon Kendra (enterprise search)
|
Azure AI Search
Hybrid retrieval: dense vector + BM25 merged via RRF; Semantic Ranker (cross-encoder reranking); integrated with Azure OpenAI for native RAG; Skillsets for NLP enrichment in indexer; Integrated vectorization
|
Vertex AI Search
Managed RAG (Vertex AI RAG Engine GA 2025); AlloyDB with pgvector; Vertex AI Feature Store; Matching Engine (ScaNN for ANN at scale); grounding via Google Search or corpus
|
pgvector + pgvectorscale LIVE
NewsRx: pgvector on Render PostgreSQL (zero incremental cost); DiskANN index; Qdrant (mid-term: SIMD Rust, 8ms p50, HNSW filtered ANN); Weaviate considered (BM25+vector RRF); Pinecone rejected (cost model)
|
| Document ExtractionOCR, forms, unstructured → structured |
Amazon Textract
OCR + ML extraction from PDFs, images; Forms (key-value pairs); Tables; Queries API (targeted extraction); Signatures; Expense and Identity docs; Lending AI (specialized); async for large batches
|
Document Intelligence
Formerly Form Recognizer; prebuilt models (invoices, receipts, ID, health insurance cards, tax W2); custom models trained on your layouts; confidence scores + bounding boxes; JSON output; batch analysis
|
Document AI
Processor library (Invoice, Contract, Expense, Identity, Lending, Healthcare); Enterprise DocAI for custom processors; Workbench for labeling; output: structured Document proto with entity extraction
|
PubMed XML + custom Python LIVE
NewsRx: XML fetch from PubMed NCBI API → Python parsing → bronze layer; Optinosis: EHR structured + imaging multimodal. No managed OCR service — purpose-built ingestion for known structured sources
|
| Pipeline OrchestrationDAGs, asset lineage, scheduling |
AWS Step Functions + MWAA
Step Functions: serverless workflow (States Language JSON/YAML); MWAA: managed Apache Airflow 2.x; SageMaker Pipelines (ML-specific DAGs); EventBridge for scheduling; Glue Workflows for ETL
|
Azure Data Factory + AML Pipelines
ADF: ETL/ELT pipelines, 90+ connectors, visual designer; Azure ML Pipelines: ML-specific steps in Python SDK; Azure Databricks Workflows; Logic Apps for event-driven; Managed Airflow in ADF
|
Cloud Composer + Vertex Pipelines
Cloud Composer: managed Airflow 2.x (GKE-based); Vertex AI Pipelines: Kubeflow Pipelines v2 SDK; native artifact lineage; KFP components; Cloud Scheduler for cron; Dataflow for streaming
|
Dagster + Modal LIVE
NewsRx: Dagster (software-defined assets, asset lineage, Components GA Oct 2025); self-hosted on Render worker; Modal handles serverless GPU steps; Prefect considered (alternative); Airflow explicitly rejected
|
| Experiment Tracking / Model RegistryRuns, metrics, artifacts |
SageMaker Experiments + Model Registry
SageMaker Experiments: run tracking, metrics, parameters, artifacts; Model Registry: versioned models, approval workflow, cross-account; SageMaker Model Cards for governance; S3 for artifact storage
|
Azure ML Experiments + Registry
Native MLflow integration (log runs, metrics, params); Azure ML Model Registry (versioned, with tags, stage transitions); Responsible AI dashboard per model; environment and dataset versioning
|
Vertex AI Experiments + Model Registry
Vertex AI Experiments: TensorBoard-backed; Metadata API for lineage; Vertex AI Model Registry: versioned, with labels and deployment history; Vertex AI Evaluation for systematic model comparison
|
MLflow LIVE
Both NewsRx and Optinosis: MLflow for run tracking, metric logging, artifact storage, model registry; runcard documentation system for IEC 62304 / 21 CFR Part 11 traceability; 27 documented prompt iterations at NewsRx
|
| Monitoring / Model ObservabilityDrift, performance, data quality |
SageMaker Model Monitor + CloudWatch
Model Monitor: data quality, model quality, bias drift, feature attribution drift; CloudWatch dashboards + alarms; SageMaker Clarify for explainability and bias; real-time or scheduled monitoring jobs
|
Azure ML Model Monitoring + Azure Monitor
AML monitoring: data drift, prediction drift, feature distribution tracking; Azure Monitor for infra metrics + alerts; Application Insights for app telemetry; Responsible AI dashboard for fairness/explainability
|
Vertex AI Model Monitoring + Cloud Monitoring
Feature skew and drift detection; prediction skew vs. training data; email/PagerDuty alerts on threshold breach; Cloud Monitoring for infra; Vertex AI Evaluation for systematic evals
|
HHEM + BERTScore + PSI LIVE
NewsRx: HHEM-2.1-Open (hallucination scoring, target ≥0.90), BERTScore F1 (target ≥0.90), Flesch-Kincaid via textstat; custom PSI monitoring for distributional shift; Dagster observability for pipeline health; Streamlit SME review interface
|
| Data Catalog / GovernanceLineage, discovery, compliance |
AWS Glue Data Catalog + Lake Formation
Glue Data Catalog: Hive-compatible metadata store, 1M+ tables; Lake Formation: column/row-level security, data sharing, governed tables; Macie for PII detection in S3; DataZone for data marketplace
|
Microsoft Purview + Unity Catalog
Purview: data map, lineage, classification, sensitivity labels, policy; Databricks Unity Catalog (Azure Databricks): column-level security, row filters, audit log; Azure Data Lake RBAC via POSIX ACLs
|
Dataplex + BigQuery Data Catalog
Dataplex: unified data governance across lakes, warehouses, marts; auto data discovery + lineage; data quality rules as code; BigQuery Data Catalog: tag templates, policy tags, column-level security
|
Alation + GxP/ALCOA+ Framework LIVE
J&J MedTech (Bishop): Alation semantic layer, data dictionary, lineage tracking, FDA+EU MDR dual compliance; Optinosis: MLflow lineage, IEC 62304, 21 CFR Part 11, HIPAA, ALCOA+ principles; C2PA + PROV-O at NewsRx; 301 governed artifacts
|
| Storage / Data LayerObject, warehouse, OLTP |
S3 + Redshift + RDS
S3: object storage (data lake foundation); Redshift: columnar OLAP; RDS PostgreSQL/Aurora; Glue for ETL; S3 + Delta Lake / Iceberg for lakehouse; DynamoDB for NoSQL; ElastiCache for caching
|
Azure Blob + Azure SQL + Synapse
Blob Storage: ADLS Gen2 for data lake; Azure SQL Database; Synapse Analytics: unified analytics (Spark + SQL pools + Pipelines); Cosmos DB (NoSQL); Cache for Redis; Delta Lake native in Synapse
|
GCS + BigQuery + Cloud SQL
GCS: object storage; BigQuery: serverless columnar analytics (pay per query); Cloud SQL (PostgreSQL/MySQL); Spanner (globally distributed OLTP); Bigtable (NoSQL at scale); Memorystore for caching
|
PostgreSQL + Medallion Architecture LIVE
NewsRx: PostgreSQL on Render; Bronze/Silver/Gold medallion layers; idempotent append-only ingestion; feed delivery audit logging; train/eval split enforced at pipeline level; pgvector on same instance (zero added cost)
|
| Serverless / Scale-to-Zero ComputeEvent-driven, no idle cost |
AWS Lambda + Fargate + Batch
Lambda: 15-min max, 10GB memory, 6 vCPU; Fargate: containerized, no server management; Batch: managed job queues for HPC/ML; EventBridge triggers; Bedrock serverless inference eliminates GPU management
|
Azure Functions + Container Apps
Functions: event-driven, consumption plan (scale-to-zero); Container Apps: serverless containers with KEDA autoscaling; Azure ML Serverless Endpoints (scale-to-zero per request); Logic Apps for workflow automation
|
Cloud Functions + Cloud Run
Functions gen2: 60-min timeout, 16GB RAM; Cloud Run: containerized, concurrency-based autoscaling, scale-to-zero; Vertex AI Serverless Prediction; Cloud Run GPUs (NVIDIA L4/T4, 2024)
|
Modal LIVE
NewsRx: Modal serverless GPU functions; per-second billing; scales to zero between runs; ~$5-50/day depending on model; $30/month free tier covers dev; native Python; H100/A100/T4 available; async batch support
|
| Fine-Tuning InfrastructurePEFT, LoRA, DPO, adapters |
SageMaker HyperPod + JumpStart
HyperPod: resilient distributed training clusters; JumpStart: 1-click fine-tune (LoRA, QLoRA) on Llama, Falcon, etc.; SageMaker Training: bring own PEFT script; S3 for dataset + checkpoints
|
Azure ML Fine-Tuning + AI Studio
Azure AI Studio: fine-tune Llama, Phi-3, Mistral with LoRA; managed compute; AML custom training with Axolotl/Unsloth in custom containers; RAFT (fine-tune on synthetic RAG data) supported
|
Vertex AI Supervised Tuning + Model Garden
Gemini fine-tuning (supervised + RLHF); open model fine-tuning via Model Garden (Llama); Vertex AI Training custom jobs (bring Axolotl container); TPU pods for large-scale fine-tuning
|
Axolotl + Unsloth + LoRA/DPO LIVE
Optinosis: Axolotl (production; multi-GPU; DPO native; YAML config); Unsloth (iteration + validation; faster, single GPU); LoRA (10-100MB adapters); QLoRA (large models on 24GB GPU); DPO (40% compute savings vs. RLHF)
|
| UI / Application LayerFront-end for AI apps |
Amplify + Bedrock Agents
AWS Amplify: full-stack React/Next.js hosting; Bedrock Agents: conversational AI with action groups; Bedrock Knowledge Bases for RAG-backed chat; AppSync (GraphQL API); CloudFront CDN
|
Azure Static Web Apps + Bot Service
Static Web Apps: React/Next.js with serverless APIs; Azure Bot Service: multi-channel conversational AI; Azure AI Studio Playground; Power Apps for low-code; Teams integration for enterprise chat
|
Firebase + Vertex AI Agent Builder
Firebase Hosting + Cloud Functions; Vertex AI Agent Builder (formerly Dialogflow CX): conversational agents; Gemini in Workspace; Looker Studio for BI dashboards; Cloud Run for custom web apps
|
Streamlit + Render LIVE
Both NewsRx and Optinosis: Streamlit for SME review UI, analytics dashboards, investor/client demos; hosted on Render; Pydantic v2 for schema validation; MCP (Model Context Protocol) for pipeline automation at Optinosis
|