Mental Model
Carton streaming is the real-time decision problem of "which orders/cartons to pick, when, at which station, in what sequence" — running every few seconds in a goods-to-person warehouse.
Think: job-shop scheduling × vehicle routing × carrier cut-time deadline constraint. The AI doesn't just optimize throughput — it satisfies hard constraints first, then maximizes efficiency within them.
Bishop's analog: Nuclear pharmacy PET isotope delivery. 2-hour isotope windows = carrier cut times. Same constraint structure, different domain.
Key Vocabulary
| Term | Definition |
| Goods-to-Person (G2P) | Robots bring product totes to human associates at fixed stations. Associates don't walk. |
| Carton streaming | Continuous sub-decision: which cartons to activate next, at which station. |
| Carrier cut time | Fixed time UPS/FedEx trucks depart dock. Hard constraint — miss it, the customer gets a "lead behind." |
| Lead behind | Home Depot term for a missed carrier cut time. North-star metric to protect. |
| Tote | Physical container robots transport. May hold SKUs for multiple orders. |
| Pick wave | Traditional WMS batch of orders. AI-native WES replaces wave planning with continuous streaming. |
| Station assignment | Routing a batch to the optimal picking station given current load. |
| Congestion | Robot/conveyor gridlock when zones are over-scheduled. The optimizer must predict and avoid it. |
| SIMPL Automation | Goods-to-person vendor HD recently acquired. Vertical lift modules, pilot at Locust Grove GA. |
Decision Architecture (Walk This Through)
Layer 1 — Eligibility Gate <100ms
Rule-based + constraint check. Which orders are eligible? (In stock, cut time not yet breached, station capacity available.) Fast filter, no ML needed here.
Layer 2 — Scoring Model (ML) 200–500ms
Supervised ML on tabular features: time-to-cut, items-per-tote, zone congestion index, station load, SKU co-location. Output: urgency-adjusted throughput score. Not an LLM.
Layer 3 — Batch Formation (Optimization) 500ms–2s
Select top-N from ranked list that minimizes robot congestion. Greedy or beam search. N = available station slots. Constraint: no zone deadlock.
Layer 4 — Station Assignment <50ms
Greedy assignment to least-loaded eligible station. Simple, auditable, reversible.
Layer 5 — Feedback Loop (Continuous) async
Actual pick time vs predicted → retrain scoring model. Congestion outcomes → update zone model. Cut-time compliance rate → rollback trigger if below threshold.
ML vs LLM Decision Rule
| Use Case | Right Tool | Why |
| Hot path (every few seconds) | Supervised ML | Structured inputs, sub-second latency, interpretable, auditable |
| Exception handling (associate reports problem verbally) | LLM | Unstructured text → structured action |
| Demand forecasting (batch, daily) | ML / time series | Tabular + temporal patterns |
| Long-horizon planning (future) | RL / simulation | Only after reward function is stable and environment well-characterized |
Interviewer trap: Defaulting to LLMs for the hot path is a weak signal. Say "supervised ML" first, always.
Validation Approach
Phase 1 — Shadow Mode
- Run AI in parallel with rules-based system — don't act on AI decisions
- Measure: would AI have decided differently? If yes, would it have been better?
- Define "better" before running: throughput, cut-time compliance %, robot travel distance
Phase 2 — Controlled A/B
- Small exposure (5–10% of decisions) → measure same metrics
- Set rollback triggers: if cut-time compliance drops below threshold → revert
Phase 3 — Full Rollout
- Only after A/B demonstrates improvement across all key metrics
- Keep rules-based as fallback hot path — not deleted, just dormant
Bishop analog: NewsRx 27-iteration controlled experimentation. Same discipline: measure before you act.
Hard Constraint Priority Order
- 1st: Cut-time compliance — no "lead behind." Non-negotiable.
- 2nd: Avoid robot/conveyor deadlock — operations stop.
- 3rd: Throughput — maximize orders per hour within constraints 1–2.
- 4th: Associate productivity — minimize idle time at stations.
Interview Language (Use Verbatim)
"The carrier cut time is the hard constraint that everything else is optimized around."
"Carton streaming has a few-second latency budget — that eliminates LLMs from the hot path."
"I decompose it into four layers: eligibility gate, scoring model, batch formation, station assignment — each with its own latency budget."
"Shadow mode before any production decision — measure the counterfactual before you trust the model."
"'Lead behind' rate is my north star. If that metric moves, everything else is secondary."
"I'd start with supervised ML with tabular features, build the feedback loop, and let RL emerge only when the reward function is well-characterized."