A Plan Is Not a Path

The gap between LLM plans and executable opportunities

What LLM Plans
Leave Out

LLM plans can look complete yet fail when future support is mistaken for money available today.

Where this question began

I am a Chinese student from a low-income family. In the second half of 2025, I used Gemini 2.5 Pro, one of the leading large language models (LLMs) at the time, to help with Cambridge outreach and application materials. I received a one-year Cambridge visiting invitation and a China Scholarship Council (CSC) award of £1,000 per month.

During visiting-student visa preparation, I opened a dedicated bank account for financial evidence. My whole family pooled money, barely enough. My grandfather then became seriously ill with cancer. The money went to his treatment, and I did not make the visit.

Qwen3.5-4BAt RMB 50,000 in initial cash

Promises feasibility
21 / 24
Completes the year
0 / 24

What do LLM plans leave out when they describe an opportunity without tracking the conditions needed to carry it out?

Follow the question

The planning blind spot

The LLM helped me navigate a demanding application process. My experience made me attentive to something a polished plan can leave implicit: the family’s ability to keep resources available throughout the journey. I wanted to test whether LLMs maintain these execution conditions as carefully as they assemble the steps. Families with little reserve have less room to absorb such omissions.

I studied one concrete condition: money available when a bill comes due. In a separate synthetic setting, Qwen3.5-4B generated full-year plans with eligibility and sufficient total annual support already secured. I tested whether those plans respected the timing of payments and receipts when promising feasibility. My experience with Gemini motivated the question; the experiment evaluates Qwen and does not simulate illness or visa financial-evidence requirements.

One opportunity, different amounts of available cash

I used one fixed Qwen3.5-4B checkpoint and twelve financial profiles. Each profile received twenty-one starting-cash amounts and two seeds: 504 responses, twenty-four at each cash level. Within a profile, the opportunity, covered tuition, costs and total award stayed fixed. Under the stated rules, arrival documents released 80% of living support; departure documents released the final 20%.

The model submitted a complete plan and judged its feasibility. I checked the choices, prerequisites and dates, then calculated whether actual receipts could pay every bill. I also searched the permitted routes to establish whether any complete route was affordable.

An award is a future resource. A payment requires spendable money on that day.

Experiment · Resource-constrained, long-horizon planning

From a proposed plan to an executable sequence

Evaluated checkpoint

Qwen3.5-4BBF16 · thinking on

Remote inference

RTX PRO 6000Blackwell Server Edition · Molab

Measured generation

504 responses · 3.63 hours6,528,329 output tokens
  1. 01 / Specify the planning problem

    A Chinese-language prompt supplies six office notices: permitted actions, dates, dependencies, costs and payment conditions. The task spans 900 days and includes a full year abroad. Four route families combine paid or waived language certification with travel advances or reimbursement. There is no unlisted borrowing or family top-up. Within each profile, only initial cash B changes.

  2. 02 / Generate the plan and its commitment

    One native-thinking response per seed and budget: a short explanation, a Boolean feasibility_claim, and an ordered list of action IDs and integer days. The claim applies to the submitted plan through the visit and final settlements. The model receives no executor feedback or repair turn. The saved final answer is evaluated separately from its thinking text.

    feasibility_claim: boolean
    actions: [{ id: string, day: integer }]
    reason: string
  3. 03 / Execute the submitted sequence

    Python first checks required stages, allowed dates, dependencies and mutually exclusive choices. It then replays dated receipts and payments, stopping before the first unpaid bill. An award becomes spendable only when its release conditions have been completed and its receipt date arrives. This separates procedural validity from financial execution.

  4. 04 / Search the reference alternatives

    A separate finite search enumerates all four supplied route families, all permitted dates and all within-day action orders. For every valid schedule it computes the minimum initial cash. The smallest is R*. Comparing B with R* distinguishes an unaffordable task from a model failure to find an affordable plan. This reference is exact within the supplied action space.

R(π)=max(0,−mink∑j≤kΔcj)R*=minπ∈ΠR(π)

For a valid schedule π, R(π) is the largest cumulative deficit across its ordered events; Δc is a receipt (positive) or payment (negative). A nonnegative year-end total cannot substitute for a nonnegative balance at every event.

Three measurements, kept separate

Positive claim: a readable true judgment. Valid sequence: all noncash constraints pass. Full execution: the plan completes every required action and payment. Cash false assurance is their decisive mismatch: a positive claim and a valid sequence, but an unpaid bill. It occurred in 97 responses; 26 other positive responses had structural errors and are counted separately.

Inspect the exact experimental setup

BF16, language-model-only inference; native thinking enabled; batch size 12. Context cap 131,072 tokens; output cap 32,768 tokens, including thinking. Temperature 1.0, top-p 0.95, top-k 20, min-p 0, presence penalty 1.5, repetition penalty 1.0. Seeds: 1729 and 2718. All 504 responses ended normally, with no output-cap truncation; 498 contained a readable Boolean judgment.

Twelve synthetic financial variants share one dependency structure. Each is evaluated at 21 common cash levels from RMB20,000 to RMB120,000, with RMB2,500 spacing between RMB35,000 and RMB70,000: 12 × 21 × 2 = 504. Monthly living costs are RMB14,500–16,500. Annual support is paid 80% after arrival confirmation and 20% after departure settlement. All twelve monthly bills remain due on their specified dates.

Qwen/Qwen3.5-4B
revision: 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
vLLM 0.26.0 · PyTorch 2.11.0 · Transformers 5.18.0
CUDA build 13.0 · NVIDIA driver 595.71.05
GPU memory reported by PyTorch: 101,975,851,008 bytes
gpu_memory_utilization: 0.85
VLLM_USE_FLASHINFER_SAMPLER=0
generation_seconds: 13,084.0199
model_load_seconds: 48.2098

The scientific figures use descriptive 95% intervals from 5,000 bootstrap resamples of the twelve whole profiles, keeping each profile’s budgets and seeds together (seed 20261006). The payment-timing diagnostic replays the original valid plans without new model calls. These are controlled synthetic results, not an estimated error rate for a population of students.

Exact input, output and implementation

Follow the same
twenty-four answers.

At RMB 50,000 in initial cash

Loading saved responses…

22 of 24 pass the checks for steps, procedures and dates.

One circle = one response, not one student.
Each stage checks the same answers separately.

Select a circle to inspect its saved result.

Twenty-two valid sequences

22 plans met the choices, prerequisites and date rules.

Twenty-one promises

With RMB50,000 at the start, 21 of 24 responses said the year was feasible.

No complete year

None paid every required bill. Valid steps still left a payment without enough available cash.

The promise arrived before the cash

Across the study, 311 answers made unconditional positive claims. Ninety-seven had valid choices, prerequisites and dates but failed a payment. Every one of those ninety-seven came from an environment with no affordable full-year route under the original rules.

The weakness concentrated where cash could fund arrival but could not sustain the year. The model correctly rejected all 142 readable answers below the arrival requirement. Where a complete route existed, 194 of 214 responses executed. The concern is a confident promise across the intervening gap.

Promises feasibilityCompletes the yearValid steps and datesReference route feasible
Promises and completed plans across initial cash
RMB 50,000

RMB 50,000: 21 of 24 promise feasibility; 0 complete the year.

All 24 answers at each saved amount. Points use actual cash spacing; lines connect existing results and do not smooth them. Download the scientific figure with descriptive intervals (English PDF).

A bill the plan could not pay

One actual response, profile 01 at RMB45,000 with seed 1729, claimed:

“The living-award installments suffice for eleven months of living costs and the later deposit; cash flow has no interruption.”

Translated from the model’s Chinese response. This is an actual model answer, separate from the reference trajectories below.

Its sequence was valid. On day 702, 330 days after arrival, a RMB14,500 living payment met only RMB3,700 in available cash. The shortfall was RMB10,800. A future entitlement could not pay that bill.

The same money, available sooner

For a given route, the cash needed at the start is its largest cumulative gap between required payments and money already received. Eligibility alone does not close that gap. Across these twelve profiles, arrival required RMB39,700–43,900; the full year required RMB50,400–60,200. These are scenario thresholds.

I then used a saved replay that moved the same final 20% to the first 80% receipt date. Every original action, date, cost and funding total stayed fixed. Ninety-five of the ninety-seven cash failures behind positive promises disappeared. Two remained because a bill still preceded the first receipt. Among all 348 valid original plans, completion rose from 194 to 305. This replay changed an institutional payment rule; it did not improve the model or generate new answers.

Follow the saved reference schedules below. The annual living award is RMB 180,000; a monthly living payment is RMB 14,500. These paths are separate from the actual model answer.

Initial cash
Same award, different receipt date

Reference profile 01 · RMB 45,000 · Original 80 / 20

Original releaseEarlier release
Follow the saved reference payments

Loading the saved reference events…

With RMB 45,000 and the original release rule, the reference stops at the first unpaid bill on day 702. Bringing the same final 20% forward makes this reference schedule complete.

Scenario days are relative, not calendar dates. The balance changes only at saved events. An unpaid bill is never deducted, and a failed path stops before that payment. Download the complete timing figure (English PDF).

Why the terms of support shape an opportunity

An invitation and an award can establish eligibility and eventual support while leaving the interval before payment unfunded. Lochner and Monge-Naranjo’s model explains how education investment depends on borrowing arrangements, including the limits imposed by public lending and private lenders.[1] Applied here, a future scholarship cannot automatically finance today’s bill. An assistant needs to understand the institution’s payment terms and the family’s available cash, alongside the award’s total.

Sen’s capability approach asks what people are actually able to do with their resources.[2] This gives the payment question its human importance: the same nominal scholarship can create different opportunities for families with different reserves. Eligibility, a promised award and a visit a student can carry out are distinct achievements. An assistant that dismisses a student on the basis of a low balance can close an opportunity before asking whether the institution can make its support available in time.

Cash also serves purposes beyond education. Deaton models assets as a buffer against income risk when borrowing is restricted.[3] That helps explain why a balance cannot simply be assigned to a visit. This suggests respecting family-care reserves in planning, alongside the family’s support for a student’s education. A student should decide which money the family has offered and which reserves protect other responsibilities.

Moynihan, Herd and Harvey describe the learning, compliance and psychological costs involved in accessing public support.[4] This framework draws attention to the work between an entitlement and its release: finding the rule, completing documents and obtaining confirmation. Applied to scholarships, assistance can reduce that work, while institutions can change the procedures and waiting periods that require families to finance access. Students with little reserve have less room to absorb these delays, even when their entitlement is secure.

There is reason to believe practical help can matter. In the randomized H&R Block FAFSA experiment, application assistance combined with personalized aid information improved applications, aid receipt and college attendance; information alone did not improve outcomes in that setting.[5] The result ties access to practical support through a difficult process. It supports a useful ambition: help students complete the steps that make support usable, and identify when the terms themselves must change.

Question 3 · Proposed improvements

Help that makes an opportunity usable

A better LLM planner must track resource conditions through every step, identify the first unmet condition, and help secure a concrete change when a path cannot be completed under current terms. The proposals below have not been tested. They aim to reduce both false assurance and false exclusion, keeping affordable opportunities open while making each plan’s dependencies explicit.

01Keep low-cash possibilities in the data

Students should receive a different answer when a relevant payment condition changes. Curate dated, institution-confirmed terms into examples that separate eligibility, promised support and money available to spend. Pair the same academic profile, costs and total award with earlier or later receipts; vary initial cash separately. Include affordable low-cash cases with explicitly available direct payment or advances, cases with no current route, and cases with missing information. Preserve all twelve living payments, deposit occupation and return, and the documents needed to release support. Label each payment state and whether any permitted complete route exists. This teaches sensitivity to a condition that matters, rather than an association between poverty and rejection. CARL studies constraint sensitivity by comparing outputs with and without constraints.[6] Here, a timing change would earn credit only when the calculated consequence and advice are correct.

02Teach the moment a payment becomes possible

A student needs to see why a plan can be carried out. For the existing 4B checkpoint, start with supervised fine-tuning through a small LoRA adapter. Train short, structured targets: required actions, receipts actually released, cash after each affordable payment, the first shortfall’s date and amount, and the cash required for a complete route. Research on process reward models shows why an apparently correct final answer can conceal faulty intermediate reasoning.[7] The adaptation here is supervision of externally checkable payment states. Add paired preference training only after this baseline: favor an accurate, useful conditional answer over both an unfounded promise and an unnecessary rejection. Numerical states and the final judgment must agree. The student-facing explanation can then name the decisive bill and the condition that would make it payable, in ordinary language.

03Turn a shortfall into a concrete request

At inference, the model would translate documents into dated events with amounts, prerequisites, release conditions and source passages. A deterministic accounting function would check each candidate, while a bounded search would examine permitted alternatives. The model would explain the result and help prepare the next action. Formal-verification research provides a precedent for using language models with constraint solvers to detect infeasibility and suggest changes.[8] In all ninety-seven cash failures here, no original route was affordable: rearranging allowed steps could not create one. Useful help must therefore surface an outside change. For each stated option—an advance from the existing award, direct payment to a provider, or an approved deferral—calculate a sufficient amount and latest usable date, then check the whole year again. Present it as conditional until the institution confirms it. Missing release terms call for a focused clarification, with the unresolved condition visible to the student.

Institutions would share the work of making the opportunity usable. A funding office could confirm one set of documents, arrange direct-paid essentials or a guaranteed short advance with clear terms, and give the student a dependable receipt date. Earlier payment keeps the nominal award fixed but transfers financing and settlement risk to the institution; it is not costless. No proposal should presume extra family money or unsafe debt. A student may choose a protected family-care reserve, and the system should compare options using only the cash they make available. The aim is a concrete, respectful conversation about what the institution can change.

How I would implement and evaluate itA small-model study with two-sided error measurement

A feasible first implementation. Keep the pinned Qwen3.5-4B checkpoint, use BF16 and one rank-8 or rank-16 LoRA adapter, and supervise short event records plus concise explanations. For a valid schedule p, let Pp(t) be cumulative obligations and Ap(t) conditional receipts already released. R(p)=maxt[Pp(t)−Ap(t)]+R*=minpR(p)The global minimum requires an exhaustive search of known, finite options. Incomplete rules or unfinished search require an unresolved, conditional answer. Projected deficits diagnose failure; actual execution still stops before an unpaid bill. An exact accounting function supplies training labels and inference feedback; no learned reward model is needed initially.

Separate missing evidence from known failure. Task-dependent calibration research distinguishes reasoning uncertainty from uncertainty in locating evidence.[9] Keep interpretation uncertainty visible even when the arithmetic is exact. Evaluate chosen reserves and hypothetical delays without inventing illness probabilities.

Evaluate both harms. Compare the base model, base plus accounting, adapter plus accounting, and an optional preference-trained adapter. Hold out whole profiles and funding-rule families; test unseen delays, equal-total timing pairs, currency rescaling and equivalent wording. Small-model planning research reports transfer and representation failures, so training success alone is insufficient.[10] Report false assurance, rejection despite a verified feasible route, payment-state error, first-shortfall accuracy and correctness of conditional changes. Ask affected students to assess usefulness and the family burden of proposed options.

What I want help to make possible

My family’s pooled money ultimately paid for my grandfather’s care. I want that decision to remain an act of care, with the student’s ambition still taken seriously. Accurate assistance can show what is possible under today’s terms, what needs an institution’s commitment, and what that commitment would have to change.

My hope is for scholarships and LLM assistance that make opportunities usable while leaving students room to care for the people they love.

Keep the evidence open to inspection.

One pinned model, twelve synthetic profiles, two seeds and all twenty-one cash amounts. These results describe this grid; they are not a real-student failure rate or a poverty threshold.

Inspect every answer All 504 saved results

Each cell is one output. Select or focus a cell for its result; arrow keys move through a panel. The columns are equally spaced categories, unlike the actual cash spacing in the curve.

CompletedCash false assurancePositive, invalid structureNegativeUnreadable interface

Loading all saved results…

The matrix retains negative answers, high-cash failures and four executable plans that the model judged false. A negative answer is not automatically an error just because a reference route exists.

Download the full scientific outcome matrix (English PDF)
How the study was checked Methods and limits

This page uses saved results from one pinned Qwen3.5-4B checkpoint (revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a), with native thinking and BF16 inference.[11] Twelve synthetic profiles were evaluated at twenty-one common starting-cash levels, RMB20,000–120,000, using seeds 1729 and 2718. All 504 outputs remain in the data. Each budget has twenty-four responses. Curves preserve actual budget spacing; the output matrix uses discrete columns.

A deterministic checker evaluated submitted actions, prerequisites, dates and receipts. A separate reference search established whether any permitted full-year route could execute. Cash false assurance required an unconditional positive claim, a valid sequence and a failed payment. Of 498 readable Boolean answers, 311 were positive; six outputs lacked a readable judgment. Four negative judgments whose own plans executed are retained, alongside structural errors and higher-cash exceptions.

The timing diagnostic reused original valid plans and moved only the final 20% to the original 80% receipt date. No new model responses were generated for that change. A failed attempted bill was not deducted from the balance. The four profile-01 reference trajectories are not model-generated plans.

This synthetic-grid study describes one model and twelve financial profiles. It estimates neither a real-student error rate nor a universal poverty threshold. The illness and financial-evidence conditions from my experience were not simulated, and the synthetic annual award is separate from my historical CSC award. The proposed improvements above are separate from these completed measurements.

The theoretical framing interprets one financial condition, not a student’s complete capabilities. Deaton models income risk; extending the insight to family care is this essay’s interpretation. The FAFSA experiment tested human assistance in a U.S. programme, rather than an LLM or changed scholarship release dates.

This is confirmation after disclosed adaptive development, not an unchanged replication. The report preserves the earlier task, clarifications, synthetic-unit confirmation and pilot. Auxiliary final-text review by six GPT-6.1-sol MAX agents interprets commitment wording; original scores and plans are unchanged. This is not human gold-standard annotation.

The scientific figure uses existing descriptive intervals that resample twelve whole profiles while retaining their budgets and seeds. They do not treat the 504 outputs as independent observations. All scientific calculations and replays preceded this essay; page controls only look up saved results.

Read and download the research English report, data and three figures
References and source contextNumbered primary sources, cost scales and evaluation background

The academic works support the interpretations and design mechanisms stated above; their authors did not test or endorse this proposal. Cost scales and distinct payment examples were checked for the report on 1 October 2026.[12][13][14] They inform the synthetic scenario and do not reconstruct my historical funding conditions. The report also draws on long-horizon planning and constraint-sensitive scheduling as evaluation background.[15][16]

  1. Lance J. Lochner and Alexander Monge-Naranjo (2011). The Nature of Credit Constraints and Human Capital. American Economic Review 101(6), 2487–2529. Publisher DOI↩
  2. Amartya Sen (1980). Equality of What?. The Tanner Lectures on Human Values, vol. 1, 197–220; lecture delivered 22 May 1979. Original lecture (English PDF)↩
  3. Angus Deaton (1991). Saving and Liquidity Constraints. Econometrica 59(5), 1221–1248. Publisher DOI↩
  4. Donald Moynihan, Pamela Herd and Hope Harvey (2015). Administrative Burden: Learning, Psychological, and Compliance Costs in Citizen-State Interactions. Journal of Public Administration Research and Theory 25(1), 43–69. Publisher DOI↩
  5. Eric P. Bettinger, Bridget Terry Long, Philip Oreopoulos and Lisa Sanbonmatsu (2012). The Role of Application Assistance and Information in College Decisions: Results from the H&R Block FAFSA Experiment. Quarterly Journal of Economics 127(3), 1205–1242. Publisher DOI↩
  6. Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang and Qiang Zhu (2026). CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs. Findings of ACL. Official proceedings DOI↩
  7. Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou and Junyang Lin (2025). The Lessons of Developing Process Reward Models in Mathematical Reasoning. Findings of ACL. Official proceedings DOI↩
  8. Yilun Hao, Yongchao Chen, Yang Zhang and Chuchu Fan (2025). Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools. NAACL. Official proceedings DOI↩
  9. Chaeyun Jang, Moonseok Choi, Yegon Kim, Seungyoo Lee, Juho Lee and Hyungi Lee (2026). Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMs. ICML; PMLR 306:50548–50571. Official proceedings↩
  10. Valerio Belcamino, Nicholas Attolino, Alessio Capitanelli and Fulvio Mastrogiovanni (2026). On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL. arXiv preprint; conference acceptance not verified. Official preprint↩
  11. Qwen Team (2026). Qwen3.5-4B model card and evaluated checkpoint. Revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. Pinned official model repository↩
  12. University of Cambridge (2026). Maintenance: postgraduate living costs. Official institutional page; checked for the report on 1 October 2026. Official cost reference↩
  13. University College London (2026). Turing Scheme at UCL. Official institutional page; checked for the report on 1 October 2026. Official payment reference↩
  14. University of Manchester (2026). Turing Scheme. Official institutional page; checked for the report on 1 October 2026. Official payment reference↩
  15. Yinger Zhang, Shutong Jiang, Renhao Li, Jianhong Tu, Yang Su, Lianghao Deng, Xudong Guo, Chenxu Lv and Junyang Lin (2026). DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints. arXiv preprint. Official preprint↩
  16. Shrenil Shaun Sharma and Avi Sharma (2026). SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling. arXiv version 2, updated 26 August 2026; preprint version cited. Official preprint↩