How to Estimate Jira Story Points with AI (and Why You Shouldn't Trust It Blindly)

7 min read
Share:

Why AI is good at this — but not good enough to skip humans

Story point estimation is fundamentally a calibration problem: given a description, predict how long it will take a typical engineer on this team to ship. AI models are surprisingly good at this for tickets that resemble past tickets. They are bad at this for genuinely novel work.

The trick is making the AI say it doesn’t know instead of confidently guessing.

Setup in Agentopias by CynetIQ

  • Go to /dashboard/refinement → Settings.
  • Pick the scale: Fibonacci (1, 2, 3, 5, 8, 13, 21), T-shirt (XS-XXL), or hours.
  • Set "Big task threshold". Default is 13 — anything the AI thinks is bigger than this gets flagged "Needs human breakdown" instead of a number.
  • (Jira) Set the custom field for Story Points (default customfield_10016).
  • (Azure DevOps) Field is Microsoft.VSTS.Scheduling.StoryPoints, no config needed.
  • The prompt that handles novel tickets

    The default refinement prompt has explicit guardrails. Here is the relevant excerpt:

    > If this ticket describes work that has no analogue in the codebase, or would require introducing a new framework / library / external service, or would touch more than ~8 files, output Needs human breakdown for the story points field instead of a number. It is much more useful to know we don’t know than to receive a confidently wrong estimate.

    We tested this. Without the guardrail, the AI gives a confident "5" for tickets like "Add a graph database for the new recommendations engine". With the guardrail, it correctly outputs "Needs human breakdown" for those tickets.

    Calibration: feeding history

    The PM agent reads up to 50 prior estimated tasks per workspace. It compares the new ticket against historic ones by:

  • File-path similarity: which files are likely touched, and what was the historic estimate for similar files?
  • Type cluster: bug fixes cluster together, new endpoints cluster together, refactors cluster together.
  • Description embedding: cosine similarity between the new description and historic ones.
  • This gives the AI a working baseline. If your team’s historic estimates are wildly inconsistent (3 sometimes means 1 day, sometimes 1 week), the AI inherits that inconsistency. Run a calibration meeting first.

    Accuracy benchmarks

    We measured against 320 estimated-and-shipped tickets across 4 teams:

    Team typeWithin ±1 pointWithin ±2 points
    Mature backend team (consistent history)84%98%
    Greenfield team (sparse history)52%81%
    Mixed bug + feature backlog76%95%
    Heavy refactor backlog68%91%
    Greenfield teams should treat AI estimates as starting points only.

    When to override

    • Cross-team dependency — AI doesn’t see the org chart. If the ticket needs another team’s API, bump the estimate.
    • Tech debt smell — AI doesn’t weigh "we should clean this up while we’re in here." Humans do.
    • Sensitive code paths — payment, auth, multi-tenant. Always require human review before the AI estimate is final.

    How Agentopias by CynetIQ writes back to Jira

    When AI Refinement runs on a Jira-sourced task, Agentopias by CynetIQ:

  • Writes the estimated points to the Story Points custom field.
  • Posts a comment on the Jira issue with the full refinement output (description, AC, points, risks, suggested assignee).
  • Logs the run on /dashboard/refinement/runs with input/output, model, cost.
  • You can roll back any refinement that landed wrong — the field history is preserved.

    Related reading

    Share:

    Agentic AI'ı denemek ister misiniz?

    Ücretsiz başlayın ve Agentopias by CynetIQ'nın 3D agentlarının geliştirme iş akışınızı yönetmesine izin verin.