In Brief
Large language models (LLMs) — including so-called “reasoning” models — plateau at 55-60% constraint satisfaction on physical constrained optimization tasks, regardless of model size or architecture. Supervised fine-tuning (SFT) improves response formatting but does not cross this threshold. Only reinforcement learning, via a method called GRPO, produces measurable gains on feasibility. This result defines a concrete boundary: current LLMs cannot replace a formal solver on problems where every constraint must be satisfied.
The Test: Electrical Flow Optimization
To measure whether an LLM can reason under constraints, you need a problem where constraints are non-optional — where a solution that violates them is simply invalid. Optimal Power Flow (OPF) constitutes this type of benchmark.
OPF is a fundamental problem in electrical networks. Given a network of nodes (generators, consumers, lines) and a power demand, the problem is to find the combination of production that minimizes cost while satisfying strict physical constraints: Kirchhoff’s laws (conservation of current and voltage at each node), generator power limits, and line transmission capacities. These constraints are not preferences — they express the physics of the network. A solution that ignores them would produce a blackout or a fire.
What makes OPF difficult for an LLM is precisely what makes it interesting: producing an “approximately correct” answer is not enough. Every constraint must be verifiable. A formal solver like Ipopt (Interior Point OPTimizer) — an open-source mathematical optimizer designed for non-linear problems — guarantees a feasible solution or declares infeasibility. It does not hallucinate.
Bernier, Ghamizi, Dogoulis, and Cordy (University of Luxembourg / LIST, 2026) built a benchmark around OPF to evaluate several frontier LLMs on their ability to reason and optimize under these physical constraints. The constraint satisfaction rate is the central metric: what proportion of the physical constraints of the problem does the produced solution satisfy?
In short: OPF is not a “find the best answer” benchmark — it is a “find an answer that breaks nothing” benchmark. A classical solver always answers correctly, or refuses to answer. An LLM, on the other hand, always answers — including when its answer is physically infeasible. That is exactly what the test measures.
The Ceiling — 55-60% and No Further
The central result of Bernier et al. (arXiv:2603.23004) is clear: across nearly all tasks requiring genuine optimization under constraints, models stagnate around 55-60% constraint satisfaction. This ceiling is independent of model architecture, size, and training regime.
What makes this result notable is its universality. So-called “reasoning” models — those that generate a long intermediate chain of thought before answering — do not obtain systematically higher scores than their non-extended-reasoning counterparts. Generating more intermediate tokens does not solve the feasibility problem.
The ConstraintBench benchmark (Xu et al., arXiv:2602.22465, 2026) confirms this diagnosis on a broader scope. This benchmark evaluates six frontier models on 200 tasks spanning ten operations research domains — production planning, vehicle routing, resource assignment, scheduling — with solutions verified by the Gurobi solver (a reference commercial solver for mixed-integer optimization problems). Models do not respond by generating code: they must produce direct decisions, verified constraint by constraint.
ConstraintBench results refine the picture:
- The best model achieves 65.0% feasibility — slightly above the OPF ceiling, but still far from a solver.
- When a model produces a feasible solution, that solution achieves 89-96% of the Gurobi optimum — meaning optimality is not the primary problem.
- No model exceeds 30.5% on the joint criterion (feasibility + optimality within 0.1% of the solver reference).
- Difficulty varies strongly by domain: 83.3% feasibility on production mix problems, versus 0.8% on crew assignment problems.
The conclusion of ConstraintBench is clear: feasibility is the primary bottleneck, not optimality. When models manage to produce a solution that satisfies constraints, it is generally close to the optimum. The problem is getting there in the first place.
In short: the 55-60% ceiling is not a random failure that more data could fix. It is a structural wall. When an LLM produces one feasible solution, it is almost optimal (89-96%) — the problem is not quality, it is reliability. Four times out of ten, the proposed solution violates a physical constraint.
Failure modes are systematic, not random. Studies identify several recurring patterns: underestimation of duration constraints (structural parsing error, not isolated), hallucination of entities absent from the problem, and a decoupling between feasibility and optimality — in some domains, models achieve high feasibility but zero optimality, because the exhaustive branch-and-bound search required for optimality is incompatible with autoregressive token-by-token generation.
SFT Learns Format, Not Reasoning
The natural response to this ceiling is to train models more extensively on solving examples. That is what Supervised Fine-Tuning (SFT) does: the model learns to reproduce correct solutions from a labeled corpus.
Bernier et al. tested this approach extensively. Their result is unambiguous: SFT improves response formatting but fails to improve physical feasibility. The constraint satisfaction rate remains at the same plateau after fine-tuning. The model learns to present its answers in the correct format — it structures its output better, names variables correctly — but it does not solve the underlying problem more effectively. In particular, it does not transfer to new network topologies (structural changes not seen during training).
This result fits within a more general observation from the 2025-2026 literature. Guo et al. (arXiv:2602.01058) show that SFT optimized in isolation can lock models into rigid imitative modes: the model learns to reproduce the patterns of the training distribution rather than generalize the reasoning. The authors’ formulation is direct — “SFT can lead to pattern matching that occasionally suppresses pre-trained reasoning logic or fails on out-of-distribution tasks.”
The distinction is important for understanding why SFT is insufficient here. OPF is not a pattern recognition problem. A slightly different network topology changes the structure of the problem. A model that has memorized solutions for known topologies does not have the tools to reason about a novel topology — because SFT did not teach it to reason; it taught it to imitate.
This mechanism also explains the results from ProOPF (arXiv:2602.03070, 2026), a professional benchmark on 12,000 OPF instances. Frontier LLMs achieve over 90% on existing benchmarks — but drop below 30% on ProOPF-B, an expert-annotated evaluation on realistic problems. The gap between performance on standard benchmarks and performance on professional problems reflects exactly the same phenomenon: apparent generalization, limited real-world feasibility.
GRPO: The Only Method That Breaks the Ceiling
If SFT does not cross the ceiling, the question is: what does?
Bernier et al. trained variants of their models with GRPO (Group Relative Policy Optimization), a reinforcement learning (RL) method. The authors report significant gains on certain network topologies — gains that SFT alone did not produce.
GRPO was introduced by DeepSeek AI in DeepSeekMath (Shao et al., arXiv:2402.03300, 2024) for mathematical reasoning. Its principle differs from supervised fine-tuning on a decisive point: instead of learning to imitate correct solutions, the model learns to optimize a reward signal. This signal directly evaluates the quality of the produced solution — here, structural validity and physical feasibility.
The mechanics of GRPO deserve explanation. In classical reinforcement learning with human feedback (RLHF), training requires a separate critic model to evaluate the outputs of the main model. GRPO eliminates this need: for each problem, the model generates a group of responses (hence “Group”), and the reward for each response is computed relative to the group average. Responses better than average are reinforced; those below are penalized. There is no external critic model — the group itself serves as the reference.
This architecture has a practical advantage: it is less expensive to train than full RLHF. More importantly, it has an advantage for constrained problems: the reward can be defined directly on verifiable criteria — are Kirchhoff’s laws satisfied? are power limits within bounds? — rather than on subjective human judgment.
In short: SFT teaches the model to imitate (“here is how a solution is presented”), GRPO teaches the model to succeed (“here is a function that says if your solution is valid, optimize it”). On problems where success is mechanically measurable (physical laws, unit tests), GRPO breaks the ceiling. On problems where “success” is subjective, the benefit fades.
GRPO results on mathematical benchmarks give an order of magnitude of the effect. On the MATH benchmark, DeepSeekMath goes from 46.8% (Instruct version, trained by SFT) to 51.7% with GRPO — a gain of five points on a pure reasoning task. On GSM8K, the gain is from 82.9% to 88.2%. These figures come from the foundational paper (arXiv:2402.03300) and are verified.
In the domain of physical constraints, the gains reported by Bernier et al. are qualitative in accessible sources — “modest but significant improvements” — without precise figures extracted from the article body. The essential finding is directional: GRPO produces an improvement in feasibility where SFT remained flat.
A recent extension, Scaf-GRPO (arXiv:2510.19807, 2025), introduces structured guidance (scaffolding) during RL training, and reports a +44.3% gain on the AIME24 mathematical reasoning benchmark compared to vanilla GRPO. These developments suggest that reward design and RL signal structure are active levers — but their systematic application to physical constrained optimization problems remains an open research area.
Implications for Practitioners
These results have concrete implications for anyone considering using LLMs in workflows requiring constrained optimization.
LLMs do not replace formal solvers. A solver like Gurobi or Ipopt guarantees feasibility or declares infeasibility. An LLM, even a high-performing one, produces solutions that violate constraints in 35-45% of cases on real problems. On problems where a violated constraint has operational consequences — electrical networks, crew scheduling, routing with strict time windows — this failure rate is unacceptable.
The distinction between code generation and direct reasoning is crucial. OR-LLM-Agent (Zhang et al., arXiv:2503.10009, 2025) — an agent based on DeepSeek-R1 and GPT-o3-mini — achieves 85% accuracy on 83 real-world operations research problems. But it does not solve constraints directly: it decomposes the problem, generates Gurobi code, and delegates optimization to the solver. This is an architecturally different pattern — the LLM acts as a problem-to-code translator, not as an optimizer. This approach works, but it circumvents the reasoning problem rather than solving it.
SFT on domain examples improves presentation, not robustness. If your pipeline includes supervised fine-tuning to adapt an LLM to optimization problems, expect the model to produce better-formatted outputs — but do not count on improvement in constraint satisfaction on out-of-distribution problems. For that, you need an explicit reward signal (RL) or a downstream solver that validates and corrects.
GRPO opens a path, but with limits to clarify. The gains reported on OPF are promising, but their precise magnitude and generalization to other domains are not yet established in the accessible literature. Reward design — what granularity, per-constraint or per-topology — remains an open question. Before considering RL training for a specific use case, you must be able to define an automatically verifiable reward: this is feasible on problems with formal constraints, harder on problems where feasibility itself is ambiguous.
Difficulty varies by domain. ConstraintBench shows a range of 0.8% to 83.3% by problem type. Production mix (a relatively well-constrained and structured task) is much more accessible than crew assignment (NP-hard combinatorics with nested constraints). Before deploying an LLM on an optimization task, evaluate precisely on representative instances — general benchmarks do not predict performance on specific topologies.
Recommended architecture by constrained profile
| Context | Recommendation | Why |
|---|---|---|
| Hard-constraint optimization (electrical grid, industrial scheduling) | Formal solver (Gurobi, Ipopt) as priority | Mathematical guarantee of feasibility or infeasibility, indispensable when violating a constraint = failure. |
| Decomposition + Gurobi code generation | OR-LLM-Agent pattern | LLM translates problem → code, solver resolves. 85% precision on 83 real OR problems. |
| Soft constraints + measurable reward | LLM fine-tuned with GRPO | Improves feasibility beyond SFT alone, via verifiable reward signal. |
| Fast heuristic + no guarantee required | Zero-shot LLM | Fast, no solver infra. Acceptable only if user validates each solution. |
Key Takeaways
-
LLMs plateau at 55-60% constraint satisfaction on physical constrained optimization tasks (OPF, operations research), regardless of model size and architecture. This ceiling is documented by at least two independent benchmarks (Bernier et al. 2026, ConstraintBench 2026).
-
SFT improves formatting of responses but does not cross this ceiling. Supervised fine-tuned models remain at the same feasibility level and do not transfer to new topologies — because they imitate patterns rather than learning to reason under constraints.
-
GRPO (Group Relative Policy Optimization), an RL method without a separate critic model, is the only approach reporting measurable feasibility gains in this domain. The mechanism: directly rewarding verifiable physical properties (structural validity, bounds compliance) rather than reproducing example solutions.
-
The real bottleneck is feasibility, not optimality: when a model produces a feasible solution, it is generally close to the optimum (89-96% according to ConstraintBench). The problem is producing a feasible solution in the first place.
-
In practice: for workflows requiring reliable optimization under hard constraints, current LLMs operate better as code generators delegating to a formal solver than as direct optimizers.