Calculator

Some of a PDF converts to audio cleanly. Some of it disappears without warning

Prose reads aloud well. A table does not, a formula does not, and a figure with its caption becomes a caption with nothing attached. The problem is that nothing announces the loss, so a document that converted badly sounds exactly like one that converted well.

DownloadApp StoreSoonGoogle Play
Free to startNo adsTR & EN

This page is about which parts survive. WeSolve+ is best for someone who wants the argument of a chapter in their ears rather than a machine reading forty pages at them.

WeSolve+ reads the whole document and writes the questions for you

Upload your PDF, photograph your notebook, or point the camera. WeSolve+ writes questions from that material, explains why each answer is right, reads the chapter back to you as a podcast, and remembers every item you missed until you own it.

Start free!

The tool below is a small browser-only tool and it is not WeSolve+: paste a few lines and text rules turn them into cards on the spot. The real app, the one that uses AI, is behind the link above.

Text to a listenable script

This is a browser-only tool, and that is all it isIt splits the text you paste by rule, and nothing else. WeSolve+ is a different thing entirely: it reads your whole PDF with AI, writes the reasoning behind every question, speaks the chapter back to you, and remembers what you missed so it can return it. Try the real app now, free!

This tool uses text rules, not AI, and it does not read files or produce sound. It takes text you have already pulled out of a document, keeps the sentences carrying the most repeated ideas, and leaves them in their original order and wording. Speaking them is your device's job, and your text stays in the browser.

First the file has to give up its text

A PDF stores instructions for drawing a page, not a stream of sentences, so a document either carries a text layer or it does not. If it was exported from a word processor it usually does. If it is a scan or a photograph of a page, the characters have to be recovered by optical character recognition, which introduces its own errors and is silent about them. The quick test takes five seconds: try to select a sentence with your cursor. If you cannot, there is no text layer and everything downstream is a reconstruction.

Then the reading order has to be recovered

Two columns, a running header, a page number, a figure caption and a footnote are all separate flows on the same page, and nothing in the file necessarily says which order a human reads them in. Working that out is layout analysis, and when it fails you hear it immediately: a sentence stops mid clause and a page number is read aloud in the middle of it. Files exported with structure, what the standard calls a tagged PDF, carry the reading order explicitly and convert far better, which is also why they work with screen readers.

What is lost silently, and this is the important part

Three things vanish without any signal. A table becomes a stream of numbers with no structure, so a row and a column sound identical. A formula becomes either nothing or a mangled phrase, unless the document carries it as MathML rather than as an image. And a figure becomes its caption, which usually says see figure four. Because none of these produce an error, a badly converted document sounds like a well converted one, and you only find out in an exam. The rule that follows is to read anything visual and listen to anything argued.

Accessibility work already solved most of this

Everything above has been studied for decades under accessibility rather than under study tools, because a blind reader has needed a document read aloud since long before anybody wanted a podcast of their lecture notes. That literature is where the answers are: tagged structure, a text alternative for every image, tables marked up as tables. It also gives you a free implementation, since every current operating system will speak selected text through built in speech synthesis, and a screen reader handles a well tagged document better than most conversion services.

A forty page document should not become forty pages of audio

Reading a chapter aloud in full takes about four times as long as reading it, and produces something nobody finishes. The useful artefact is short: the argument, the two or three anchor facts, and the thing you keep forgetting. That is what the tool above builds, by keeping roughly one sentence in five of whatever you paste, in your own wording. Two minutes of audio holds about three hundred words, and a recap you can get through on the walk to a lecture is worth more than a complete recording you never open. There is more on when listening beats reading on the podcast generator page.

The recap the app makes from the file itself

Upload the PDF, photograph the page or take a camera shot and WeSolve+ reads it and produces a spoken recap of roughly one minute, along with an explained quiz and a card deck from the same file. The recap is deliberately short for the reason given above, and it is built from the document rather than from your extracted text, so you do not have to do the pulling out yourself. On the free plan there are two recaps a week. At wesolveapp.com in any browser and on the App Store for iPhone and iPad. Figures on pricing, and the feature on audio recap.

Sources used on this page

Part of a document, and whether it survives being spoken
What is on the pageConvertsWhat to do instead
Ordinary proseYes, cleanlyListen to it, ideally as a second pass
A numbered listMostlyKeep it short, numbers are hard to hold
A tableNo, structure is lost silentlyRead it, then say the conclusion aloud
A formulaNo, unless tagged as MathMLRead it and write it out by hand
A figureOnly its caption survivesLook at it, then describe it to yourself
A scanned pageOnly after character recognitionCheck a paragraph for recognition errors

Does this tool convert a PDF?

No. It works on text you have already pulled out. A browser rule cannot open a file, and this page exists partly to explain why that step is harder than it looks.

How do I know if my PDF has a text layer?

Try to select a sentence with your cursor. If you cannot, the page is an image and the characters have to be recovered by optical recognition, with its own errors.

Why do tables sound wrong?

Because speech has no columns. A table read aloud is a stream of numbers where a row and a column are indistinguishable, and nothing warns you that the structure was lost.

What is the best way to get audio from a document?

Use the read aloud already built into your operating system on a well structured file. It handles tagged documents better than most conversion services and costs nothing.

Should I listen to a whole chapter?

Rarely. Reading it aloud in full takes about four times as long as reading it, and the recording gets abandoned. A two minute recap of the argument is the thing people actually finish.

Is anything I paste uploaded?

No. The selection runs in your browser and makes no network request. Closing the tab ends it.

Last updated: 2026-08-15