SHIPPED · v0.2.0
1LocalAITool v0.2.0 — reading pictures
New
Photos, scans and screenshots back into text you can edit. Switch on *Read pictures* under Skills and Tasks gains a task for it: drop in one picture or a whole stack — a ZIP of them, or a scanned PDF, works too — and choose what comes back: a Word document, a PDF, plain text, Markdown, a web page or JSON. Every page lands in one document in the order the pictures sort (page 2 before page 10), or tick *One file per picture* for a ZIP with a file for each.
Four readers, under This computer › Models: LightOnOCR 2, the default, which keeps tables, headings and accents as they were; PaddleOCR-VL 1.6, about three times faster when plain text is enough; and Qwen3-VL in two sizes, which follow an instruction such as “only the totals”. Each downloads once, is checked against its published checksum, and runs in its own short-lived process, so nothing stays loaded after the job.
Phone photos come out the right way up — the rotation a phone notes in the file is honoured — and WEBP, TIFF and, on a Mac, iPhone HEIC photos are read as well as JPG and PNG.
Ask about a picture in chat. Attach a photo, a receipt or a screenshot and ask what it says or shows. When the chat's model cannot see, a reader on this computer looks first — describing the picture and copying out its words — and the model answers from that. Follow-up questions about the same picture work without attaching it again. If *Read pictures* is off, the chat says so and switches it on for you in one click.
Send pictures from anywhere: the phone's queue form takes several photos at once, POST /v1/ocr takes them from a script, 1lat read ~/Scans/ -f docx saves the document beside you, and the MCP tool read_pictures lets an agent hand over a folder of scans.
PDFs keep every alphabet. A PDF now borrows a font this computer already has and embeds just the letters it uses, so Vietnamese, Greek, Cyrillic, Chinese and Japanese are drawn instead of turned into question marks, and the text can be selected and searched. Tables are drawn as tables.
Two new document skills: *Make a Word document* and *Make a text file* leave a .docx or .txt beside any answer.
Switches show the right way round. The toggles on Skills, Routines and the sidebar's Pause drew their knob from the middle, so an off switch looked on and an on switch lost its knob.
Open a document straight from the inbox. A finished job's Word file or PDF opens in its own app; *Show file* still finds it on disk.