Skip to content
StrataHub

Manufacturing · June 14, 2026 · 6 min read

AI-Enabled Last-Mile Manufacturing Resource Planning

ERP tells you what happened. The last mile of manufacturing planning — today's machine down, tomorrow's rush order — still runs on spreadsheets and gut feel. Here is how we combine LLMs with operations research to fix that.

Every manufacturer we work with has an ERP, an MRP run, and usually an APS license someone fought hard for. And yet the decisions that actually determine whether Friday's shipments go out — a press down for four hours, a rush order from the biggest customer, two operators out sick — get made on a whiteboard, in a huddle, by whoever has the most tribal knowledge.

We call this the last mile of manufacturing resource planning: the gap between the weekly plan the system produces and the hourly reality on the floor. It is where most of the money leaks out, and it is exactly where the current generation of AI finally earns its keep.

Why the last mile stays manual

Classical planning systems assume the inputs are right and the world holds still. Neither is true. Setup times in the routing master were entered in 2019. Yields vary by shift. The MRP run finished at 2 a.m. and was stale by 7.

The deeper problem is interaction cost. Re-planning in an APS means someone who knows the tool rebuilding a scenario, which takes 45 minutes. So the plant doesn't re-plan; it improvises. The plan and the floor diverge a little more every hour, and by Thursday the schedule is fiction.

LLMs are the interface, not the optimizer

Here is the architecture pattern we deploy, and the division of labor matters enormously.

The optimizer is classical operations research: a MILP or constraint-programming model (we typically use OR-Tools or Gurobi) that encodes machines, labor, changeover matrices, material availability, and due dates. It is deterministic, auditable, and fast — a plant-scale scheduling model re-solves in seconds to a few minutes.

The LLM sits in front of it as a conversational layer. A production supervisor types, or says: "Line 3 is down until 2 p.m. and the Meridian order moved up to Wednesday. What breaks?" The agent translates that into parameter changes — capacity on line 3 zeroed for the morning, a due date tightened — re-runs the solver, and answers in plain language: which orders slip, by how much, and what the recovery options are. Move the changeover, authorize overtime, split the batch across lines 2 and 5.

Do not let the LLM do the math. Language models are unreliable at combinatorial optimization and will confidently produce infeasible schedules. Every number in the answer must come from the solver; the model's job is translation, explanation, and orchestration.

This split is what makes the system trustworthy. The solver guarantees feasibility. The LLM makes it accessible to someone who has never opened the APS client and never will.

Scenario planning becomes a conversation

The unlock isn't answering one question — it's the second and third question. "What if we ran a Saturday shift instead?" "What does that do to the changeover count?" "Show me the version where Meridian slips one day but nothing else moves."

Each of these is a solver run with modified constraints, compared against a baseline. In one engagement, a mid-market industrial components maker went from one formal scenario per week (built by a single planner) to 40–60 scenarios per week generated by supervisors themselves. Decisions that waited for Monday's planning meeting now get made in the moment, with the trade-offs quantified: expedite the rush order and you eat $8,400 in overtime and push two B-tier orders by a day — or decline and protect the schedule.

The planner didn't lose a job. They stopped being a human query interface and started doing actual planning.

What it takes to build this for real

Data plumbing first. The solver is only as good as its inputs: live machine status, current WIP, actual labor on shift, inventory positions. That usually means integrating MES or SCADA signals, the ERP order book, and a labor system into a near-real-time operational store. In most engagements this is 50% of the work, and skipping it produces a very articulate system that reasons about a plant that doesn't exist.

A tool-calling contract, not free-form generation. The agent operates through a fixed set of typed tools — set_capacity, update_due_date, run_schedule, compare_scenarios. Every solver invocation is logged with its full parameter set, so any recommendation can be replayed and audited weeks later.

Evals before rollout. We build a regression suite of 100+ real scenario requests harvested from planner interviews, each with a known-correct parameter translation. The agent has to hit >95% exact-match on constraint edits before a supervisor ever sees it. Paraphrase robustness matters: "line 3 is dead," "we lost the big press," and "no capacity on 3 until after lunch" must all resolve identically.

Observability in production. Every conversation, tool call, solver run, and accepted-or-rejected recommendation is traced. Acceptance rate by user and scenario type is the single most honest health metric — when it dips, either the data is stale or the model is mistranslating, and we want to know which within the day.

Where to start

Don't start with the whole plant. Pick one line or one cell with a painful, recurring re-planning problem — the packaging bottleneck, the shared CNC bank — and build the loop end to end: live data in, solver in the middle, conversation on top, decisions logged.

That is a Pilot-shaped problem: 4–6 weeks to a working system on real data that supervisors use on real disruptions. If the acceptance rate holds, Co-Build scales it across lines and adds the harder constraints — tooling, operator certifications, maintenance windows. The pattern is production-or-nothing for a reason: a scenario-planning demo on synthetic data impresses everyone and changes no decisions. A system that answered Tuesday's breakdown correctly changes how the plant runs.

Work with us

Shipping something like this?

We co-build production AI systems with enterprise teams — pilots in 4-6 weeks.