Dashboard

How to Prompt AI to Estimate a Project Timeline

A model will happily produce a project timeline from a one-line description, and that timeline will be useless. The prompt that produces a usable estimate does three things the naive version does not: it forces the model to decompose the work before assigning any numbers, it makes it state the...

Steve Jefferson
Steve Jefferson
Developer Advocate
15 August 20261 min read

A model will happily produce a project timeline from a one-line description, and that timeline will be useless. The prompt that produces a usable estimate does three things the naive version does not: it forces the model to decompose the work before assigning any numbers, it makes it state the assumptions each number depends on, and it asks for a range with a named worst case rather than a single figure. Estimation is a decomposition problem. Get the decomposition right and the arithmetic is easy.

The reason this is worth doing at all is not that a model estimates better than you. It is that it estimates differently, and the gap between its breakdown and yours is where the tasks you forgot are hiding.

Why the obvious prompt fails

Ask "how long will it take to build a booking system for a dental practice" and you get four to six weeks. You will get roughly that answer for a bike shop, a barber and a law firm, because the model is producing the median of every similar sentence it has seen rather than reasoning about your project.

Three specific failures follow:

Anchoring on a plausible number. Once the model states six weeks, everything after that is rationalised to fit six weeks, including the task list.

Silent omission. Data migration, access control, the client review cycle, staging setup and the two weeks of fixes after launch are typically missing. They are also usually 40 percent of the real timeline.

False precision. "3.5 days" for a task nobody has scoped implies a confidence that does not exist and is difficult to argue with in a client meeting.

The prompt structure that works

Four moves, in this order. The order is the technique.

1. Give it the constraints before the question

The model cannot estimate what it does not know, and it will not ask unless you make asking cheap. Front-load the facts that change the answer:

text
Context:
- Team: me (full stack) plus a designer at 2 days/week
- Stack: Next.js, Postgres, Stripe. Existing codebase, ~15k lines
- Client: 8-person dental practice, one decision maker, slow to reply
- Hard constraint: must integrate with their existing practice
  management system, which has a REST API and no sandbox
- Definition of done: staff can take bookings in production,
  patients get confirmations, no manual steps
- What already exists: auth, billing, admin shell

That "no sandbox" line will move the estimate more than anything else in the brief, and a model would never have guessed it. The wider habit is covered in giving AI context about your business.

2. Force decomposition before numbers

This is the single highest-leverage instruction in the whole prompt:

text
First, break this into tasks at the level where a task is
half a day to three days of work. Do not estimate anything yet.
Include tasks for setup, integration, testing, review cycles,
deployment and post-launch fixes, not just feature building.
List them and stop.

Ending with "list them and stop" prevents the model racing to a total. You then read the list, which is the actual work. Add what it missed, delete what does not apply, and only then move on.

3. Estimate each task with a stated assumption

text
Now estimate each task in working days as three numbers:
optimistic, likely, pessimistic. For each, state in one line
the single assumption the "likely" number depends on.
Where you have no basis for an estimate, say "unknown, needs
a spike" instead of guessing, and say what the spike would test.

Two things happen. The assumptions become a checklist you can verify, and "unknown, needs a spike" surfaces the genuinely risky items instead of hiding them inside a confident average. On a project with an undocumented third-party API, that flag will appear exactly where the schedule will actually break.

Giving the model explicit permission to say it does not know is the difference between a risk register and a fantasy. The same principle drives getting AI to ask clarifying questions.

4. Aggregate deliberately, not by adding

Do not let it sum the likely column. That produces a number that is wrong in a predictable direction, because every task's likely case assumes nothing goes wrong, and something always does.

text
Now produce:
- Sum of optimistic, sum of likely, sum of pessimistic
- A recommended commitment date using
  (optimistic + 4 x likely + pessimistic) / 6, rounded up to
  the nearest half week
- The three tasks with the widest spread between optimistic
  and pessimistic, which are where the schedule risk lives
- What would have to be true for the optimistic total to happen

That weighted formula is the standard three-point estimate. It is not magic, but it systematically pulls the answer away from the optimistic case, which is the direction human and model estimates both err in.

The final question is the one to bring to a client meeting. "The four-week version requires their IT contact to answer within a day and the practice management API to work as documented" is a sentence that reframes a deadline as a set of dependencies, which is exactly what it is.

Calibrate against work you have already done

The step almost nobody takes, and the one that turns this from a party trick into a tool.

Take three finished projects. Run the same prompt against their original briefs, with no knowledge of the outcome. Compare the model's estimate to what actually happened.

You will get a ratio, and it will be consistent. If the model lands at 0.6 of your real timelines, that multiplier is now part of your process, and it is more useful than any improvement to the prompt itself. It also tells you where the shortfall is: usually client response time, integration surprises, and the tail of small fixes after launch.

Sensible use and dishonest use

Sensible: using the decomposition to find the tasks you forgot, using the wide-spread items to decide what to spike first, using the assumption list as the basis for a conversation about what the client controls.

Not sensible: pasting the total into a fixed-price proposal. You are the one who carries the overrun, and a model that has never seen your client's approval process cannot price that risk. Use the output as input to your estimate, then commit to your own number.

Where the estimate becomes a commercial document, the framing in writing an AI project proposal matters more than the arithmetic, and the assumption list you generated is the strongest defence you will have when the scope starts moving. Handling scope creep on an AI project picks up there.

Frequently asked questions

Should I tell it the deadline I am hoping for? No. Anchor the model with a target and it will produce a plan that fits the target. Estimate first, compare to the deadline second, and let the gap be visible.

Does a bigger model give better estimates? Marginally. The prompt structure matters far more than the model, because the failure was never a reasoning failure, it was a missing-context failure. The technique in prompt engineering that applies here is decomposition, not model selection.

Can I use this for a project in a domain the model does not know? Yes, with more of your context in step one and more scepticism in step three. Expect more "unknown, needs a spike" entries, which is the correct output for an unfamiliar domain rather than a failure.

How do I handle the client who wants one number? Give them one number, the weighted commitment date, and one sentence of what it assumes. The range is your working document; the commitment is what you say out loud. Never show a three-column table to someone who asked for a date.

Once the timeline is set, the same decompose-before-you-answer approach carries into tracking the work against it; see how to prompt AI to turn meeting notes into action items.

Related reading: what is prompt chaining.

How did this land?

About the author

Steve Jefferson
Steve Jefferson

Developer Advocate

Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.