Manufacturing AI & Predictive Maintenance — Conceptual Crash Course  |  Carrier Round 1  |  2026-05-18

1. Mental Model

Predictive Maintenance (PdM) is the application of ML to equipment sensor data to detect anomalies before failure — replacing scheduled maintenance (wasteful) and reactive repair (costly). The core loop is: collect sensor signals → build health baseline → detect deviation → predict RUL → trigger intervention.

For HVAC systems like Carrier's, sensors measure temperature, pressure, power draw, vibration, and refrigerant flow. A fleet of 10,000+ units becomes a training corpus. The ML challenge is heterogeneity: equipment age, model variants, install environments. The architectural challenge is scale: streaming ingest, unified feature store, single model generalizing across all variants.

Carrier's actual implementation: AWS Glue + SageMaker. Raw IoT data → S3 lake → Glue ETL → SageMaker training on historical fault labels → real-time inference → Abound platform alerts.

4. Validation / QA Angle

John's QA background maps directly here:

  • Ground truth labels — fault labels are like test acceptance criteria. Model output must match known outcomes
  • False negative cost — missed fault = unplanned downtime or safety event (same as a missed defect in regulated systems)
  • Precision vs. recall tradeoff — same trade-off John made at Invistics: tune threshold to minimize missed diversions (false negatives), accept some false positives
  • Confidence scoring — RUL predictions need uncertainty intervals, not point estimates. Same as HHEM score bands at NewsRx
  • Human-in-the-loop — high-risk equipment flags go to technicians for confirmation before intervention (same pattern as Invistics)

2. Key Concepts

TermDefinition
Remaining Useful Life (RUL)Predicted time until equipment failure. The primary output of many PdM models.
Health IndexA normalized 0–1 score representing equipment condition. Aggregates multiple sensor signals.
FMEAFailure Mode & Effects Analysis — systematic mapping of failure modes to causes. Informs feature selection.
Condition MonitoringContinuous sensor data collection to track equipment state over time.
Anomaly DetectionUnsupervised ML approach: learn normal operating envelope, flag deviations. First layer before supervised fault classification.
Digital TwinA virtual replica of physical equipment, fed by real sensor data, used for simulation and what-if analysis.
Feature StoreCentralized repository of engineered features (time-window aggregations, ratios) for consistent model training and inference.
Drift DetectionDetecting when model performance degrades because equipment characteristics or operating conditions have shifted.
OEEOverall Equipment Effectiveness — the standard manufacturing KPI combining availability, performance, quality.

3. Manufacturing Demand Forecasting

What it is: Predicting future demand for products/parts to optimize inventory, production schedules, and supply chain capacity. For Carrier: predicting HVAC unit sales by region/season, spare parts demand, service call volume.

Architectural pattern: Time-series ML (ARIMA/Prophet for baselines; gradient boosting or LSTM for complex patterns) fed by historical sales, weather data, economic signals. Feature engineering is the critical work — holiday effects, weather correlations, supply disruption signals.

John's bridge: Cardinal Health — SAP-based order fulfillment and distribution analytics for 80+ radiopharmacies. Same problem: forecast radiopharmaceutical demand by imaging center, by isotope half-life window (2-hour delivery). The precision requirement (radioactive decay) is actually harder than HVAC.

5. AWS AI Stack for Manufacturing (Carrier's Actual Stack)

LayerAWS ServiceRole in PdM
IngestIoT Core, KinesisReal-time sensor stream from equipment
StoreS3 Data LakeRaw + processed sensor history
TransformAWS GlueETL, feature engineering, time-window aggregation
TrainSageMakerModel training, HPO, experiment tracking
ServeSageMaker EndpointsReal-time inference on new sensor readings
GenAIAmazon BedrockConversational AI ("Tell Me More" in Abound)
OrchestrateStep Functions / LambdaPipeline automation, alerting workflows

Bridge: John's Experience → Carrier's Stack

  • S3 Data Lake → J&J MedTech (Cloudera to AWS migration, S3 + Redshift)
  • SageMaker → Amgen (SageMaker in the regulatory AI stack)
  • Kinesis streaming → GeoSpatial Metrics (real-time UAV telemetry processing — same streaming architecture, different sensor type)
  • Bedrock GenAI → NewsRx (self-hosted equivalent — Qwen3-32B via vLLM; Bedrock is managed Bedrock version of same pattern)

6. Interview Language

"The architectural pattern for predictive maintenance is the same I've applied in regulated data environments — anomaly detection on time-series sensor data, federated feature stores, SageMaker-based deployment with drift detection. The domain vocabulary is different; the architecture is not."
"At Invistics, we processed ~1M transactions per day through anomaly detection models at 96% accuracy, deployed to 300+ sites — that's the same fleet-scale inference problem Carrier is solving with HVAC equipment."
"For demand forecasting, my Cardinal Health work was essentially demand prediction for radiopharmaceuticals — time-critical products with hard delivery windows — feeding SAP order fulfillment across 80+ facilities. The forecasting patterns are structurally the same."
"Digital twin architectures are a natural evolution of the canonical data models I build — the real-time feedback loop between the physical asset and its virtual representation is an MLOps pattern at its core."