Additional Q&A
Q: How would you turn a rules-based WES into AI-driven?
"I wouldn't rip and replace — I'd instrument and shadow. First, log every rule-based decision with its inputs and outputs. Second, train an ML model on that logged data — now you have a model that learns what the rules already know. Third, shadow the model against live rules to find where it diverges. Fourth, investigate the divergences: does the model find something the rules miss? Only then does the model get decision authority, and only with a rollback path back to the rules."
Q: What metrics would you optimize and why?
"In priority order: (1) Cut-time compliance rate — non-negotiable, customer impact. (2) Robot travel distance per order — proxy for congestion and efficiency. (3) Throughput — orders per hour. (4) Associate idle time — labor efficiency for hourly workers. I'd define these before writing a line of model code. Undefined success criteria is how AI projects fail."
Q: What information would you need before finalizing the design?
"Five things: (1) Current decision latency — how fast does the rules system decide today? That's my latency budget. (2) Historical event log — I need pick outcomes, cut-time compliance, and robot positions to train on. (3) Failure mode catalog — what breaks the rules system? Those are my first test cases. (4) Warehouse topology — zone map, conveyor layout, station capacity. (5) SIMPL Automation API surface — what data does the G2P system expose in real-time?"
Q: How would your system adapt to congestion or missed cut times?
"Congestion: the scoring model includes a zone congestion index as a feature — when zone X is saturated, orders requiring zone X get deprioritized. Cut time breach: trigger is a hard rule, not ML — if an order is T-minus 15 minutes from cut, it gets forced to top of queue regardless of ML score. The ML optimizes efficiency; the rules protect the hard constraints. They coexist."
Weak Spot Defense
| If asked... | Say this exactly |
| Java preferred, you're Python-primary | "I architect the decision layer in Python — the inference service is Python. The integration layer uses whatever the team owns in Java. I've integrated with Java/Spring services throughout my career." |
| No direct warehouse experience | "I frame this as problem types. Carrier cut times are 2-hour isotope delivery windows. Same constraint structure — I've solved this problem. I learn the operational domain from the people who live in it." |
| AWS vs GCP | "The architecture patterns are identical. BigQuery = Redshift, GKE = EKS, Pub/Sub = Kinesis. The GCP-specific APIs are a sprint-one learning curve, not a design blocker." |
| No RL / Computer Vision experience | "For carton streaming, RL is premature until you have stable reward signals and a well-characterized environment. I'd start with supervised ML with measurable, interpretable features and build toward RL as the system matures." |
| Ruby on Rails / web frameworks gap | "At this level I'm designing the decision layer that the web layer calls. I work at the service interface, not the controller." |
Technical Q&A
Q: How would you architect OR → QOP → Rule Based → LLM → Vector Store → Firestore?
"Each layer handles what the one below it can't. OR — pure mathematical optimization when variables are known and constraints are well-defined. QOP — queue management and priority ordering layered on top. Rule Based — business logic codified for operations to own: hard cut-time rules, line capacity limits, readable by warehouse managers. LLM Reasoning — everything rules can't pre-program: FedEx early/late, line down, Black Friday. Vector Store — the LLM's memory: historical decisions and outcomes embedded for semantic retrieval, RAG pattern, find the nearest analog to this situation. Firestore — operational state persistence, GCP-native, every agent reads and writes here, shared real-time view of the warehouse."
Q: How do you handle the explore/exploit tradeoff in a live warehouse?
"Exploit heavily during peak hours — cut times are live, no experiments. Explore during low-traffic windows using shadow decisions logged but not acted on. Never explore on orders within 30 minutes of cut time."
Q: What's your model retraining strategy?
"Triggered retraining when distribution shift is detected — e.g., pick time prediction error rises above threshold. Scheduled nightly retrain as baseline. No manual retraining — it's a pipeline, not a notebook."
Questions to Ask Naaga
"What does the current WES decision logic look like — pure rules, or is there any ML already in the hot path?"
"The SIMPL Automation acquisition — how does the AI layer interface with their goods-to-person system? What does the API surface look like?"
"What's the data latency from warehouse floor event to decision system today?"
"How does the team currently handle cut-time exceptions — is that a manual process or system-driven?"
"You've seen HD's technology evolution from the inside for 7+ years. What's surprised you most about what problems AI actually solves versus what you expected?"
"What would make someone fail in this role in the first 90 days?"
Say-This-Not-That
| DON'T SAY | SAY INSTEAD |
| "I'd use a reinforcement learning agent" |
"I'd start with supervised ML, build feedback loops, let RL emerge when the reward function stabilizes" |
| "LLMs can handle carton streaming" |
"Carton streaming is structured inputs/outputs — supervised ML, not LLMs. LLMs sit in the exception interface." |
| "I don't have warehouse experience" |
"I've solved this constraint type in nuclear pharmacy logistics — same problem, different domain." |
| "I know Python, not Java" |
"The inference layer I architect is Python. I design interfaces so your Java services call the decision engine cleanly." |
| "The optimal algorithm would be..." |
"Before committing to an algorithm, I'd need to see the event log, the latency budget, and the failure mode catalog." |
Agent Architecture
| Agent | Owns |
| Order | Classify order, assign carrier window priority |
| Pick Routing | Robotic vs. human pull per line item, timing |
| Consolidation | Mixed orders — coordinate timing so robotic + hand-pull arrive together |
| Wave/Batch | Group orders into batches working back from 10am/2pm windows |
| Carrier Assignment | Which orders hit 10am vs. 2pm FedEx |
| Staffing | Live station capacity; reroute on no-shows; 400 vs. 200 mode |
| Exception (LLM) | FedEx early/late, line down, Black Friday surge |
Magic Forge
Internal engineering platform running BigQuery-based monitoring (BQM) over the streaming infrastructure (Pub/Sub, Dataflow). Observability layer over live warehouse streams. Not public. HD is 100% Gemini/Google. If tooling comes up — Gemini Enterprise, Vertex AI, Google ecosystem. Claude is a test case only.
Watch-Outs
Overqualification signal: 40 years of career depth can read as "too senior." Lean in — frame as pattern recognition speed, not seniority. "I've seen this problem before in a different domain."
Travel 10–20%: They're explicit about it. Be enthusiastic: "I want to see the warehouse. You can't architect what you haven't observed — I'd be there as much as they'd have me."
Academic trap: Any answer that starts with theory instead of operational constraint is a weak signal. Always lead with the constraint, then the mechanism.
LLM reflex: Pause before answering any design question. Ask yourself: "Is this a structured decision problem?" If yes — say ML first.
GCP unfamiliarity: Don't pretend to know Vertex AI internals. Own the AWS-to-GCP translation: "Same patterns, different APIs — I'd be productive by week two."
Digital Twin Framing — Power Move
The carton streaming decision system is a digital twin by definition — a real-time virtual replica of warehouse state (robot positions, station loads, conveyor congestion, inventory, order queue) queried to evaluate decisions before committing to physical action.
Bishop's resume already has this: DARPA digital twin simulations (FORTRAN/Cray). The bridge is direct and credentialed.
Deploy if: interviewer mentions "simulation," "what-if modeling," "warehouse digital twin," or asks how the system evaluates decisions without acting on them. Say: "What you're describing is the pattern I know as a digital twin — I architected this exact capability for DARPA. The warehouse state model is the twin; the scoring model queries it before any robot moves."
SOTA Quick-Reference
| Item | Fact |
| Cloud | GCP only. 10-yr partnership. |
| Data WH | BigQuery (15+ PB) |
| ML Platform | Vertex AI |
| Robotics | SIMPL Automation (acquired, G2P) |
| Pilot site | Locust Grove, GA (near Atlanta) |
| Backend lang | Java Spring Boot (primary) |
| Streaming | Cloud Pub/Sub + Dataflow |
| GenAI product | Magic Apron (customers), Gemini Enterprise (associates) |
| Legacy tech | Informix 4GL, Hadoop (still in ops) |
| Phase | Build — 1-2 AI people today |
Bishop's Strongest Analog
Nuclear pharmacy PET isotope delivery = carrier cut-time problem. 2-hour isotope half-life windows = UPS truck departures. Both: hard constraint, time-critical, customer consequence if missed. Cardinal Health: 80+ radiopharmacies, hub-and-spoke, 2-hour delivery to 90% of US imaging. Say this early.