Markdown camera

Image to Markdown

A photographed page comes back as an editable, block-structured note — not a picture you still have to type out. Photos and handwriting use Moondream; only a server-confirmed dense screenshot switches to 2–3 Llama 4 Scout parts.

or drop JPEG / PNG / WebP files here

Only photos you explicitly confirm are uploaded · paid account required · originals and results are not stored automatically

Three steps from photo to note

  1. Snap or pick a page

    “Take a photo” opens the rear camera on a phone; on a desktop, choose or drop up to 3 JPEG, PNG, or WebP images.

  2. Check, order, confirm

    Photos land on the workbench shown below — compressed locally, free to reorder or rotate. Nothing uploads until you press “Convert to Markdown”.

  3. Review flagged words, export

    The image to Markdown result comes back block by block. Returned confidence is preserved and low values are highlighted; missing confidence is shown as unavailable. Edit, copy, or export as Obsidian or plain .md.

Image to Markdown workbench screenshot with the page stack, the original handwritten photo, and the Markdown result column awaiting confirmation
The actual workbench after one handwritten page is added — recognition starts only after you confirm.

Keep this result by copying or exporting it

Process

From one photo to editable knowledge

  1. Capture up to three pages

    Use the rear camera on a phone or pick files on a desktop. Pages can be reordered, rotated, and removed before anything uploads.

  2. Confirm the upload

    Photos are first oriented, downscaled, and compressed in your browser. The page tells you exactly how many compressed copies will be sent before you convert.

  3. The server selects a bounded recognition path

    Photos and handwriting stay on Moondream. Only when server-side pixel analysis confirms a dense screenshot is it split into 2–3 overlapping parts for Llama 4 Scout. Exact boundary overlap is removed; uncertain text is kept for review.

  4. You review the uncertain parts

    When the selected model returns numeric confidence, values below the threshold are highlighted next to the original. If confidence is omitted, the UI says so instead of inventing a score.

Use cases

Made for pages that never had a text layer

A PDF exported from Word carries its text with it; a photograph of your own handwriting carries nothing but pixels. Turning that image to Markdown is a different job from PDF conversion, and this page is built specifically for it:

Lecture and reading notes

Handwritten notes become searchable Markdown you can revise and link into a vault; headings and bullets are preserved when the model recognizes them.

Whiteboards and meeting pages

Photograph the board before it is wiped. Recognized checkboxes can become Markdown task items, ready for you to verify before pasting elsewhere.

Printed handouts and forms

Worksheets, one-page briefs, and annotated printouts that only exist on paper convert the same way — the model reads print more reliably than handwriting.

Privacy

What uploads, and what never does

The boundaries differ by tool: Word .docx conversion stays in the browser; PDFs are analyzed locally and then use Cloudflare conversion or Moondream OCR. In the Markdown camera, photos and handwriting use Moondream; only a dense screenshot confirmed by the server is split into 2–3 temporary parts and sent to Llama 4 Scout.

What is sent is the compressed working copy of the photos you confirmed — nothing is scanned from your camera roll, and originals stay on your device. The copies are processed for recognition and are not kept as your files; results remain only in the current session until you copy or export them and are not saved to your account automatically.

  • Upload happens only after you press convert, never on selection.
  • Compressed copies are sent; the full-resolution originals never leave the device.
  • Temporary screenshot parts do not become extra pages: ordering and page points still follow the one original image.
  • Recognition requests are not used to build a library of your images — there is no server-side storage of results today.
Limits

What it reads well, and what it still gets wrong

An honest image to Markdown converter should say where it fails. Clear handwriting and print are easier inputs, and the model can return headings, lists, checkboxes, and simple tables. Every result still needs judgement:

  • Needs your judgement

    Very loose or overlapping handwriting — inspect the source even when the model omits confidence.

  • Needs your judgement

    Dense tables with drawn lines and merged cells; complex chemical or matrix notation.

  • Needs your judgement

    A dense screenshot may cross a crop boundary. The merge removes only exact overlap, so uncertain boundary text is retained and still needs review.

  • Needs your judgement

    Mind-map style pages where arrows carry the meaning. Marks are described in the output, but the drawing itself is not reproduced.

  • Needs your judgement

    Low light, strong shadows, or steep angles. Re-shooting flat and bright helps more than any model setting.

A vision model can still produce plausible mistakes. Treat low or missing confidence as a review signal and compare important text with the photo.

Workflow

Where the Markdown goes next

Drop-in exports

The output is ordinary Markdown, so it lands anywhere Markdown does. The Obsidian export adds frontmatter and a heading outline so a captured page files straight into a vault; the plain export pastes cleanly into GitHub, Notion, or a docs site.

Questions

Image to Markdown FAQ

Is this image to Markdown conversion free?

Live image recognition requires a paid account and costs four page points for each successfully recognized original. A server-confirmed dense screenshot may use 2–3 temporary parts, but it remains one logical page and is charged once. Failed conversions cost zero points.

Are my photos uploaded or stored?

Photos are uploaded only after you sign in and explicitly confirm, as compressed working copies, for recognition only. Originals never leave your device, and the service does not keep a store of your images or results. Results are not saved automatically, so copy or export them before leaving.

Can it read handwriting, and in which languages?

Legible Chinese and English handwriting is transcribed with Moondream, and print is usually easier. The source language is preserved. Confidence is shown only when the model returns it; important text still needs review against the photo.

How many pages can I convert at once?

Up to three original images per conversion. Order them in the page stack before converting; Markdown follows that order. A dense screenshot can be split into 2–3 recognition parts only after server confirmation, but those parts remain one logical page.

Why not just photograph notes into a generic AI chat?

You can, but you re-type the prompt every time and the output format drifts. Here the pipeline is fixed: validated blocks, deterministic Markdown rendering, explicit missing evidence, and one-tap Obsidian export.

What about HEIC photos from an iPhone?

Photos taken directly with the in-page camera arrive as JPEG. If you pick an existing HEIC photo and your browser cannot decode it, the page says so immediately — re-exporting as JPEG (or taking a screenshot of the photo) solves it.

Photograph one page and check the result

Use a paid account for live recognition, review the Markdown beside the source photo, then copy or export it before leaving.