All BuzPulse apps
AI & Insights

Smart File Extractor

Pull the text and data out of a screenshot, PDF, Word file, spreadsheet or deck.

Extract text and data from the files you already have. Drop in a screenshot, PDF, Word document, spreadsheet, presentation, CSV or text file and get clean, editable content back. Documents are read entirely in your browser and never uploaded — a PDF, .docx, .xlsx, .pptx or .csv is parsed on your own machine, so nothing is stored anywhere. Screenshots are read by AI: you get the text, plus the links, email addresses, phone numbers, dates, amounts and any table rebuilt into rows you can paste into a spreadsheet. Error screenshots are explained in plain language. Nothing is guessed — a value that is not legible is left out rather than invented.

What Smart File Extractor gives you

Everything below is included — no add-ons, no per-app checkout.

  • Screenshots and images read by AI — text, links, numbers and tables
  • PDF, Word, Excel, PowerPoint, CSV and text read in your browser
  • Documents are never uploaded — parsed on your own machine
  • Spreadsheets and tables come back as rows you can paste
  • Error screenshots explained in plain language, with things to try
  • Nothing invented: unreadable values are left out

About Smart File Extractor

What it does, how to use it, and where it stops.

Text locked inside a file cannot be searched, edited or pasted anywhere else. Smart File Extractor takes the files you already have — a screenshot, a PDF, a Word document, a spreadsheet, a presentation, a CSV or a plain text file — and hands back the words and the numbers as content you can actually use. Images are read by AI; every other format is parsed in your own browser and never uploaded.

Extract text from images and screenshots

Paste a screenshot straight from the clipboard with Ctrl+V, or ⌘V on a Mac — most people already have it there, so that is usually the whole job. You can also drag an image in or pick one from your device. PNG, JPG, WebP, GIF and BMP are accepted, up to 6 MB.

Press Smart Extract Everything and you get the text back along with whatever else was in the picture: links, email addresses, phone numbers, dates and amounts, and a table rebuilt into rows if the image genuinely contained one. Narrower buttons — Text Only, Table, Links and Translate — are there for when you want one part of the answer rather than all of it.

Images are the one format read by AI rather than parsed locally, because a picture has no text layer to read. There are only pixels, and something has to recognise the letter shapes in them.

PDF text extraction

A PDF made by a computer — exported from Word, generated by an accounting system, downloaded as a bank statement — carries a real text layer underneath the page. That layer is read directly in your browser, page by page and in order, for up to the first 100 pages.

Nothing is uploaded and no AI allowance is spent, so a long PDF costs you a few seconds and nothing else. What comes back is a continuous document you can search, edit and paste, rather than a picture of a page.

Word and DOCX extraction

A .docx file is a ZIP of XML underneath, so it is unpacked and read directly on your machine. You get the paragraphs in document order, the headings Word itself marks — Title, Subtitle and the Heading styles — kept as headings rather than flattened into the body, and any tables as rows and columns you can copy.

Only the modern .docx is read. The older binary .doc is a different file format entirely and is not supported.

Excel and XLSX data extraction

Every worksheet in the workbook is read, and each one keeps its own sheet name so you can tell the figures apart. Rows come back as rows: where the first row is a complete set of column names with data beneath it, it is treated as headers — and where it is not, the data is left exactly as it is rather than a header row being invented.

The result pastes back into a spreadsheet with its columns intact, which is the whole point. A number you can total is worth more than a screenshot of one. Only .xlsx is read, not the older .xls.

CSV and TSV extraction

The separator is detected from the file rather than assumed — comma, tab or semicolon, whichever it actually uses. That matters everywhere the comma is the decimal mark and exports come out semicolon-separated, which is most of Europe and much of the Middle East.

Quoted fields containing commas or line breaks are handled properly, so an address column does not scatter itself across three cells.

PowerPoint and PPTX extraction

A deck is read slide by slide and in order, with each slide numbered so you can find your way back to it. The title placeholder on each slide is recognised as a title rather than treated as one more line of body text, and tables placed on slides are kept as rows.

It extracts what a deck says, not what it looks like. Speaker notes, images, animation and layout are not part of the result.

TXT, MD and LOG extraction

Plain text files — .txt, .md, .markdown and .log — are read exactly as written, with nothing reformatted or interpreted. Markdown stays as Markdown.

It is the plainest path through the tool, and it exists so that a log file someone sent you goes through the same window as everything else instead of needing a different one.

OCR for images and scanned PDFs

Optical character recognition is the step that turns letter shapes in a picture back into editable characters. It uses the surrounding words as context, which is why a clear photograph of ordinary language works far better than a sharp image of a random serial number — and why contrast and focus matter more than file size.

Every image goes through it, and so does one particular kind of PDF. A file produced by a scanner or a phone camera is a photograph of a page with no text layer at all, and when a PDF turns out to have no readable text the tool says so and reads its first page as an image rather than pretending the page was blank. That first page is the part that gets read: a fifty-page scan does not come back in full.

Supported file formats

Images: PNG, JPG, WebP, GIF and BMP, up to 6 MB each. Documents: PDF, DOCX, XLSX, PPTX, CSV, TSV, TXT, MD and LOG, up to 25 MB each.

The two limits differ because the two paths cost different things. An image is sent away to be read, so it is bounded by what that request can carry; a document is parsed in your own tab, where the only thing it can exhaust is your own memory. Anything outside those lists is refused with a sentence saying why, rather than accepted and failed later — video, audio, executables and ZIP archives are not processed.

Privacy and browser processing

PDFs, Word files, spreadsheets, presentations, CSVs and text files never leave your device. They are read by code running in your own tab, so there is no upload to fail, nothing on a server to delete afterwards, and no copy of your document for anyone to lose.

Images are the exception, and the tool says so on screen before you press the button: reading a picture needs a vision model, so the image is sent to BuzPulse’s AI provider for exactly as long as it takes to read it. It is not written to storage, not saved to the database, not logged, and never used for advertising. Nothing you put through this tool is kept by BuzPulse afterwards.

Where people use it

A PDF you need to quote from

Pull the wording out of a report, contract or statement so you can paste the relevant paragraph into an email instead of retyping it.

A supplier price list in Excel

Read every sheet at once and paste the rows straight into your own workbook, with the sheet names and columns intact.

An error message on screen

Capture the dialog and extract the exact text and error code so you can search for it or send it to support, rather than transcribing a long code by eye.

A deck sent instead of a document

Get what the slides actually say, numbered slide by slide, without opening PowerPoint to read them.

A receipt photographed on a phone

Pull out the merchant, date, line items and total and put them straight into your expense record.

A CSV export from another system

Check what is really in the file — separator detected, quoted fields intact — before importing it anywhere.

Tips

  • Zoom in before screenshotting. More pixels across each letter is the single biggest accuracy gain available.
  • A PDF that comes back with no text is a scan, not a broken file. That is the case OCR is for, and the tool offers it automatically.
  • Do one column at a time for multi-column images, or the reading order comes out interleaved.
  • Always read the result against the original. Digits and short codes are where errors hide.
  • Documents are free of the daily AI allowance, so a PDF, spreadsheet or deck can be run as often as you like.

What it will not do

  • Handwriting is not reliably recognised.
  • Only the first page of a scanned PDF is read as an image.
  • PDF text is read to 100 pages; anything beyond that is not included.
  • The older .doc, .xls and .ppt formats are not read — only .docx, .xlsx and .pptx.
  • Layout is not preserved. Expect the words and the numbers, not the design.
  • Video, audio, executables and archives are not accepted at all.

How it works

The things worth knowing before you sign up.

How do I extract text from a PDF?

Drop the PDF in and press Extract Content. The text layer underneath the page is read directly in your browser, page by page and in order, for up to the first 100 pages — nothing is uploaded and no AI allowance is spent. If the file turns out to be a scan with no text layer, the tool says so and reads its first page as an image instead.

How do I extract text from a Word document?

Choose the .docx file and press Extract Content. It is unpacked and read on your own machine: the paragraphs in document order, the headings Word itself marks kept as headings rather than flattened into the body, and any tables as rows and columns you can copy. Only .docx is supported — the older binary .doc is a different format and is not read.

How do I extract data from Excel?

Drop the .xlsx in and press Extract Content. Every worksheet is read and each one keeps its own sheet name, so figures from different sheets stay apart. Where the first row is a complete set of column names with data beneath it, it is treated as headers; where it is not, nothing is invented. The result pastes back into a spreadsheet with its columns intact. Only .xlsx is read, not the older .xls.

Can I extract data from a CSV or TSV file?

Yes. The separator is detected from the file rather than assumed — comma, tab or semicolon, whichever it actually uses — which matters wherever the comma is the decimal mark and exports come out semicolon-separated. Quoted fields containing commas or line breaks are handled properly, and a header row is recognised when there is one.

Can I extract text from a PowerPoint presentation?

Yes. A .pptx deck is read slide by slide and in order, with each slide numbered so you can find your way back to it, the title placeholder on each slide recognised as a title, and tables placed on slides kept as rows. It extracts what the deck says, not what it looks like: speaker notes, images and layout are not part of the result.

How do I extract text from an image?

Paste it with Ctrl+V, or ⌘V on a Mac, drag the file in, or pick it from your device, then press Smart Extract Everything. PNG, JPG, WebP, GIF and BMP are accepted, up to 6 MB. You do not have to say what is in the picture: the tool reads it, works out what it is looking at, and gives you the text along with anything else it found.

Can I extract text from a screenshot?

Yes — it is what this tool was built for, and pasting is still the fastest way to do it. Take the screenshot, press Ctrl+V here, and press Smart Extract Everything. The text comes back as normal editable text rather than a picture of text, with its own Copy button, and Download .txt saves the whole result as a file.

What file formats does Smart File Extractor support?

Images: PNG, JPG, WebP, GIF and BMP, up to 6 MB each. Documents: PDF, DOCX, XLSX, PPTX, CSV, TSV, TXT, MD and LOG, up to 25 MB each. Anything else is refused with a sentence saying why, rather than accepted and failed later — video, audio, executables and ZIP archives are not processed.

Is the file extractor free?

Yes, on the free BuzPulse plan. Documents are parsed in your own browser, so a PDF, spreadsheet, deck or CSV costs nothing at all and has no daily limit. Image extraction uses AI, and free accounts get the standard fair-use allowance of 20 actions a day shared across the BuzPulse AI tools — spent only on a successful extraction, so a failed attempt costs you nothing. Paid plans have no daily limit.

Are my documents uploaded or stored?

PDFs, Word files, spreadsheets, presentations, CSVs and text files never leave your device — they are read by code running in your own tab, so there is nothing on a server to delete afterwards. Images are the exception, and the tool says so on screen before you press the button: reading a picture needs a vision model, so it is sent to BuzPulse’s AI provider for as long as it takes to read. Nothing is saved to storage, written to the database, logged, or used for advertising.

What is OCR?

Optical character recognition — software that finds letter shapes in an image and turns them back into editable text, using the surrounding words as context to resolve ambiguous characters. Every image here goes through it, and so does a scanned PDF, which is a photograph of a page rather than a document with text in it.

Can it read a scanned PDF?

Yes, one page of it. A scanned PDF has no text layer to read, so when the tool finds none it tells you and renders the first page as an image to be read by OCR instead. That first page is what comes back — a fifty-page scan is not returned in full.

Can I extract a table from a file or screenshot?

Yes, when there really is one. Spreadsheet sheets, Word tables, CSV data and tables on slides come back as proper rows and columns, and a table in an image is rebuilt the same way. Each table has two copy options: Copy pastes straight into a spreadsheet with the columns intact, and CSV gives you comma-separated data. A bulleted list or a form is not turned into a table, because that would invent structure the file does not have.

Can I extract links, emails and phone numbers?

Yes, from images. Links come back clickable, email addresses and phone numbers as tap-to-contact, each individually copyable. Numbers and currency amounts are pulled out too, with whatever label was printed beside them, which is what makes a screenshot of a dashboard or an order confirmation useful rather than just readable.

Can I analyse an error screenshot?

Yes, and it is one of the things this tool is built for. When the picture is a software, browser, Windows, Mac or mobile error, you get a plain-language account of what the screen is saying, the exact message and any error code, and practical things to try. It will not invent a diagnosis: a screen that only says "Something went wrong" gets told back to you as exactly that, with general steps, rather than a confident guess at a cause the image cannot show.

Does it work on receipts and invoices?

Yes. A photographed receipt or invoice comes back as named fields — merchant, document number, date, the line items with quantity and price, subtotal, discount, tax, total and currency. Only the fields actually legible in the image are shown, and nothing is recalculated: if the printed total disagrees with the lines, you are shown what is printed.

What about text in another language?

For images, the language is identified and the original text is always kept exactly as it appears. Smart Extract adds an English translation alongside it when the text is not English, and there is a Translate button if that is the only thing you want. The original is never replaced by the translation.

How accurate is it?

Text read straight out of a PDF, Word file, spreadsheet or CSV is exact — it is the file’s own text, not a guess at it. Images are different: very good on clear, straight, high-contrast text, less so on angled photographs, low light or unusual fonts. Contrast and focus matter far more than file size.

What if my image is blurry?

Accuracy drops, and errors can look plausible rather than obviously wrong. Re-capture at a higher zoom if you can, and always read the result against the image.

Can it read handwriting?

Not reliably. It is built for printed and on-screen text.

Smart File Extractor is included on the free plan

$0 forever

Works well with

Included in the same account, at no extra cost.

Merge, split, rotate and convert PDFs — in your browser, never uploaded.

Documents
Open app

Shrink images, PDFs, Office files and more in your browser, with the real saving shown.

Utilities
Open app

Business writing help: emails, replies, proposals and everyday drafting.

AI & Insights
Open app

Build a clean, branded invoice and send it as a PDF in minutes.

Finance
Open app