Free · no signup · nothing stored

How good is your AI prompt?

Paste any prompt below. You'll get a score out of 100 and the specific things that are missing — measured against seven things that change what an AI gives back.

0 / 5,000
0/100
Want this fixed automatically?

Frompting rewrites your idea into a structured prompt using one of 60 frameworks.

Try it free

What prompt scoring is

Prompt scoring measures how completely a prompt specifies what you want, before you send it — then names what is missing.

It checks for the things that measurably change an answer: a clear task, an audience, an output format, constraints, an example. It does not judge whether your idea is good. It judges whether a model has enough to act on, or whether it has to guess.

Most disappointing AI output is not a model failure. It is a model filling in blanks the prompt left open — and filling them differently each time you ask.

One prompt, scored at every step

Every score below is real. Paste any row into the checker above and you will get the same number — the scoring is deterministic, so it never varies.

What changedAdded to the promptScore
Starting point Write a blog post about email marketing. 26 weak
+ audience
who it is for, and what they do not know
for small e-commerce owners who have never run a campaign 44 weak
+ format and length
the shape of the answer
900 words, four sections with headings 64 workable
+ role
who the model should be
You are a senior lifecycle marketing strategist 72 solid
+ constraints
what a good answer looks like
avoid jargon, do not invent statistics, keep the tone practical 79 solid
+ an example
show, do not only describe
Open each section with a one-line takeaway. Here is an example of the style I want: "Send the welcome email within an hour or it loses half its value. Most shops wait a day and wonder why nobody opens it." 84 solid

Seven words became eighty-one, and the score went from 26 to 84. Nothing clever was added — each step answered a question the model would otherwise have answered for you, differently every time. Note where it stops: even with an example, this prompt does not reach 100, because it never asks the model to think before drafting. For a blog post that is the right call.

The eleven parameters

Weights reflect how much each one changes an answer in practice. Every check awards partial credit: mentioning a thing earns some of the points, doing it properly earns the rest.

Clear task12

An action verb and the artefact. "Write a 600-word launch email" scores full; "write about our launch" earns half, because "about" leaves the deliverable open.

Context and audience12

Naming an audience earns part. Qualifying them earns the rest — "for finance leads who have never used a data warehouse" tells a model what to explain and what to assume.

Output format12

"Use sections" could mean two or twenty. "A three-column table comparing X, Y and Z" cannot be misread. The commonest reason a good answer arrives unusable.

Specific, not vague12

Scored on filler density. "Good", "relevant" and "engaging" do not under-specify slightly — they hand the decision to the model entirely.

Examples10

The biggest lever most prompts leave unused. One short example says more about tone and structure than a paragraph describing them.

Constraints10

Graduated by how many you set. One rule earns little; three start to define what "done" means — tone, priorities, what to leave out.

Reasoning instruction8

"Think it through step by step." It changes analysis, comparisons and decisions more than almost anything else here — and does little for short creative tasks.

Role or persona8

Weighted lower than most people expect. "You are a writer" is barely narrower than nobody; seniority and domain earn the full weight.

Explicit exclusions6

What to stay away from. Separate from constraints because it fixes a different failure: the confidently wrong inclusion nobody ruled out.

Length target6

Without a number, output length is a coin flip — and length drives depth. 400 words and 1,500 words are different answers to the same question.

Readable structure4

Applies above 60 words. A long prompt in one block buries its own instructions; one requirement per line survives.

Four patterns that reliably work

Beyond filling the eleven slots, a handful of techniques change results consistently enough to be worth learning by name.

Few-shot — show one example

Give one or two examples of the output you want before asking for a new one. It communicates tone, structure and level of detail faster than any description, and it is the technique people most often skip. Even a bad example marked "not like this" helps.

Chain of thought — ask for reasoning first

"Work through it step by step before answering." On multi-step problems — analysis, comparisons, anything with trade-offs — this reliably improves the conclusion. On a one-line creative task it adds noise, so it is not a universal upgrade.

Delimiters — fence your input

When a prompt contains material to work on — a draft, a transcript, data — wrap it in triple backticks or tags. Without a fence, models routinely confuse "the text I am giving you" with "the instructions I am following".

Role — but a specific one

"You are a senior tax accountant specialising in SaaS revenue recognition" shifts vocabulary and assumed knowledge. "You are a helpful assistant" does nothing at all: it is what the model already believes.

What actively hurts a prompt

Some habits do not merely fail to help — they make output worse.

  • Contradicting yourself. "Be comprehensive but keep it under 200 words" forces the model to pick, and you cannot predict which it drops.
  • Politeness padding. "I hope you can help me with something, if it is not too much trouble" costs input space and adds nothing. Be direct.
  • Asking for citations. A model without browsing will invent plausible-looking sources. Ask it to name well-known works instead.
  • Requesting charts or images from a text model. Ask for a markdown table.
  • Page counts. A model has no page. Use word counts.
  • Burying the ask. Three paragraphs of background before the actual request. Lead with the task, then give context.

The checker flags the fabrication-prone ones as advisories rather than deductions — they are rare in hand-written prompts, and a surprise penalty on a rare pattern reads as arbitrary.

Prompt frameworks, and when they help

A framework is a reusable slot structure — a checklist for the eleven parameters above, arranged for a particular kind of task. They matter because a blank page is the hardest place to start, not because the acronym is magic.

FrameworkStands forBest for
CO-STARContext, Objective, Style, Tone, Audience, ResponseGeneral-purpose writing where tone matters
RACERole, Action, Context, ExpectationProduct and UX tasks with a defined actor
APEAction, Purpose, ExpectationShort, clear asks where speed beats structure
TAGTask, Audience, GuardrailsBounded work where what to avoid matters
RISERequest, Input, Scenario, ExpectationDelegation and briefs handed to someone else
STARSituation, Task, Action, ResultAchievement narratives, interviews, case studies

These six are a sample. Frompting's library holds 61 across eight domains, and the generator picks one for you based on what you have written.

When a framework does not help: when you already know exactly what you want. A framework is scaffolding for an under-specified idea; if your prompt already scores in the eighties, adding one adds ceremony rather than clarity.

Scoring image and video prompts

The checker is built for text prompts, and about half of it transfers.

Transfers well: specificity, exclusions, constraints, a clear task. An image prompt lives or dies on concrete nouns, and "no text, no watermark, no extra limbs" is exactly the sort of exclusion that saves a generation.

Does not apply: length targets and reasoning instructions. An image generator has no word count and does not deliberate. Expect those to read as gaps when they are not, and ignore them.

Visual prompts also have their own concerns the checker does not model — camera and lens language, lighting, aspect ratio, negative prompts, and style references. If you are working in images, treat the score as a partial signal rather than a verdict.

The five mistakes that cost the most points

  1. No audience. The most common omission and among the most expensive. A model writing for everyone writes for no one.
  2. Format left implied. You picture a table; you get six paragraphs. You picture bullets; you get an essay. Say it.
  3. Vague quality words. "Make it engaging" is not an instruction, it is a hope. Replace it with something checkable: "open with a concrete example, no adjectives in the first sentence".
  4. No exclusions. Saying what you do not want is faster than correcting it afterwards, and it is the cheapest way to stop invented statistics and filler.
  5. No example. Skipped almost universally, worth 10 points, and usually the fastest improvement available to a prompt already scoring in the seventies.

How the scoring works

The checker is deterministic. No language model grades your prompt — it is pattern analysis in plain PHP, so the same prompt always returns the same score. That matters more than it sounds: ask an AI to rate a prompt twice and you will often get two different numbers, which makes it useless for telling whether an edit actually helped.

Nothing is stored. Your prompt is scored in memory and discarded — not logged, not saved, not used for training.

Common questions

Does a higher score mean a better answer?

It means fewer decisions are left to the model. A low-scoring prompt can still get a good answer by luck, when the model happens to guess your audience and format correctly. A high-scoring prompt gets a usable answer repeatedly.

Should every prompt score 100?

No. Some checks do not apply to some tasks — a tweet does not need a reasoning instruction, a quick factual question does not need a worked example. Anything at 70 or above gives a model what it needs. Chase the specific gaps that matter for your task rather than the number.

Why did my prompt lose points for something it has?

Each check looks for the thing done concretely. "Use a good format" mentions format without specifying one, so it earns partial credit rather than full. The card tells you which half is missing.

Does prompt length matter?

Only as a proxy. Twenty words is usually too few to carry a task plus context; past about 600 the instruction starts getting buried in background. What matters is what the words do, not how many there are — two 50-word prompts in our testing scored 43 and 100.

Do these rules apply to every AI model?

Broadly yes. Clarity, audience, format and constraints help every model. Newer reasoning models need the "think step by step" instruction less, because they do it internally, but nothing on this list makes output worse.

Glossary

  • Prompt engineering — writing inputs so a model reliably produces what you need, rather than needing several attempts.
  • Zero-shot — asking without examples. What most people do.
  • Few-shot — including one or more examples of the desired output.
  • Chain of thought — asking the model to reason before concluding.
  • System prompt — standing instructions that apply to a whole conversation, separate from any single message.
  • Temperature — how much randomness the model uses. Lower is more predictable, higher more varied.
  • Hallucination — confident output that is not true. Prompts asking for citations or statistics invite it.
  • Negative prompt — things to exclude. Standard in image generation; useful as plain exclusions in text.
  • Token — the unit models read and bill in, roughly ¾ of a word.

Skip the editing

If a prompt is scoring in the forties, closing every gap by hand takes longer than rewriting it. Frompting builds a structured prompt from a rough idea — audience, format, constraints and length already specified.

Try it free