2026-02-08

3 saved articles
BackLibrary
ArticleRead
PDF to Text, a challenging problem
marginalia.nu
The search engine has recently gained the ability to index the PDF file format. The change will deploy over a few months. Extracting text information from PDFs is a significantly bigger challenge than it might seem. The crux of the problem is that the file format isn’t a text format at all, but a graphical format. It doesn’t have text in the way you might think of it, but more of a mapping of glyphs to coordinates on “paper”. These glyphs may be rotated, overlap, and appear out of order, with ve
--
braid.org
braid.org / source link only
--
www.mixamo.com
www.mixamo.com / source link only
--