Calculator

Anyone can write questions. The hard part is deciding which ones count

Most tests fail before a single question is written, because nobody decided what the test was for. The tool below takes text you already have and hands it back as items you can answer; the page around it is about the decision the tool cannot make for you.

DownloadApp StoreSoonGoogle Play
Free to startNo adsTR & EN

WeSolve+ is best for someone building a test from material they already own, a chapter, a handout or a marking scheme, rather than from a generic bank. Paste your content, choose the output, and the questions come back with the answers hidden until you ask.

WeSolve+ reads the whole document and writes the questions for you

Upload your PDF, photograph your notebook, or point the camera. WeSolve+ writes questions from that material, explains why each answer is right, reads the chapter back to you as a podcast, and remembers every item you missed until you own it.

Start free!

The tool below is a small browser-only tool and it is not WeSolve+: paste a few lines and text rules turn them into cards on the spot. The real app, the one that uses AI, is behind the link above.

Text to test items

This is a browser-only tool, and that is all it isIt splits the text you paste by rule, and nothing else. WeSolve+ is a different thing entirely: it reads your whole PDF with AI, writes the reasoning behind every question, speaks the chapter back to you, and remembers what you missed so it can return it. Try the real app now, free!

This tool uses text rules, not AI. It keeps numbered questions intact with their options, matches an "Answer:" line to the item above it, and turns plain statements into gap fill items. It cannot judge whether an item is a good item, which is what the rest of this page is about. Your text stays in the browser and is never sent anywhere.

Decide what the test is for, and the questions follow

There are two jobs and they pull in opposite directions. A test that ranks people against each other is norm referenced and needs items that split the group, which means throwing out anything everyone gets right. A test that asks whether somebody has met a standard is criterion referenced and keeps those easy items, because passing means being able to do the thing. Most people writing a test for themselves want the second and unconsciously build the first, then feel bad about a score that was never meant to be a percentage of anything. Deciding this first is the whole of backward design: start from what a person should be able to do, then write only items that would show it.

A blueprint stops you testing the easy third

Left alone, anybody writing questions drifts toward whatever is easiest to ask, which is definitions. A blueprint fixes that by writing the test as a grid before any question exists: topics down the side, and across the top the kind of thinking each item demands, which is what Bloom's taxonomy was built to label. Decide how many items go in each cell, then fill them. Two things fall out immediately: you notice the topic you were about to skip, and you notice that eighteen of your twenty items were recall. The same grid is how curriculum mapping catches a syllabus that was taught but never assessed, and it takes about ten minutes.

Validity and reliability are two different failures

Validity asks whether the test measures the thing you meant. Reliability asks whether it would give the same answer twice. A test can be perfectly reliable and completely invalid: a twenty item test on chapter one, given to somebody revising the whole syllabus, will produce the same score every time and tell you nothing about the syllabus. The usual internal reliability figure is Cronbach's alpha, and it rises with length, which is why a five item test cannot be relied on no matter how good the items are. In practice the useful move is not calculating anything; it is asking a colleague to read the paper and say what they think it is testing.

How long, and what a short test cannot tell you

Length buys stability rather than coverage. Under about ten items, a single lucky or unlucky question moves the score by ten percentage points, so a short test can rank two people wrongly without anybody making a mistake. Twenty to thirty items is where a topic level test starts behaving. Timing is a separate decision and it changes what is measured: a generous limit tests knowledge, a tight one tests fluency, and both are legitimate as long as you know which one you chose. If the test is a rehearsal for a real exam then the real exam's timing is part of the thing being rehearsed, which is the argument for a proper mock rather than an untimed set.

Afterwards, find out which questions were broken

The scores are the least interesting output. Item analysis asks two questions of every item: how many people got it right, and did the people who did well on the test as a whole also do well on this item. An item that strong candidates get wrong while weak ones get right is not hard, it is broken, usually because the wording admits a second reading. An item everybody answers correctly is not testing anything, though on a criterion referenced test that may be exactly what you wanted. The formal machinery behind this is item response theory, but the two by two version above catches most of it, and it is the only step that makes next term's test better than this one's.

Building one from material you already have

Upload a PDF, a photograph of a page or a camera shot and WeSolve+ reads it and writes an explained test from that file, along with a deck and a short audio recap. Explained is the part that matters here: an item with reasoning attached can be checked against your blueprint, and a bare letter cannot. Items you get wrong are kept in their own list, which is the beginning of the analysis described above. It runs at wesolveapp.com in any browser and as an App Store app on iPhone and iPad, free to begin, metered by file rather than by item. Figures on pricing.

Sources used on this page

A twenty item blueprint for one chapter, filled in before writing anything
TopicRecall itemsApply or analyse items
Definitions and vocabulary40
The core process24
Worked calculation04
Common misconception12
Links to the previous chapter12

Is this test maker free?

Yes, and there is no account. The page does the work on your own machine, so there is nothing to bill and nothing to rate limit.

Does my material get uploaded?

No. It is handled by the page you are on. Load the page, disconnect, and the box still works, which would be impossible if it needed us.

How many items should a test have?

Twenty to thirty for a topic. Under ten, one lucky question moves the score by around ten points, so short tests rank people wrongly without anyone making an error.

Can it write the wrong answers for me?

Not here. A plausible wrong answer requires knowing the subject well enough to know how people go wrong, which is not something a text rule can do. The app writes those.

What is a blueprint?

A grid you fill in before writing questions: topics down the side, type of thinking across the top, and a target count in each cell. It takes ten minutes and it stops you writing eighteen definition questions.

How do I know if a question was bad?

Check whether the people who scored well overall got it right. If strong candidates miss an item that weak ones get, the item is ambiguous rather than difficult.

Last updated: 2026-08-15