<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><description>Machine Learning Librarian at @hf.co&#xA;</description><link>https://bsky.app/profile/danielvanstrien.bsky.social</link><title>@danielvanstrien.bsky.social - Daniel van Strien</title><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mwr55bl5cs2e</link><description>Uploaded a dataset of 98,877 historical newspaper pages (1700s–1940s) to the Hub, each with its original OCR text, word boxes and confidence scores. &#xA;&#xA;https://huggingface.co/datasets/biglam/europeana_newspapers_images</description><pubDate>30 Sep 2026 20:02 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mwr55bl5cs2e</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mw7aqrjxws22</link><description>The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages.&#xA;&#xA;The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dropped are back.&#xA;&#xA;huggingface.co/datasets/biglam/europeana_newspapers&#xA;https://huggingface.co/datasets/biglam/europeana_newspapers</description><pubDate>23 Sep 2026 17:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mw7aqrjxws22</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mvng7ycauc2r</link><description>Made some improvements to the OCR scripts onboarding in my uv-scripts collection.&#xA;&#xA;First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models.&#xA;https://huggingface.co/datasets/uv-scripts/ocr</description><pubDate>16 Sep 2026 15:08 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mvng7ycauc2r</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mvap3l5j3k2j</link><description>Need a historical illustration?&#xA;&#xA;Search 1.49 million images from British Library books and Britannica (1500s–1920s).&#xA;&#xA;Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.&#xA;&#xA;https://huggingface.co/spaces/davanstrien/historical-illustration-search</description><pubDate>11 Sep 2026 13:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mvap3l5j3k2j</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mvap3l5j3k2j</link><description>Need a historical illustration?&#xA;&#xA;Search 1.49 million images from British Library books and Britannica (1500s–1920s).&#xA;&#xA;Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources.&#xA;&#xA;https://huggingface.co/spaces/davanstrien/historical-illustration-search</description><pubDate>11 Sep 2026 13:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mvap3l5j3k2j</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mv5qhszdrk2i</link><description>Used Astra + Jobs to see how well this model performs on the full olmOCR-bench&#xA;&#xA;Unsurprisingly, it doesn’t do brilliantly overall: 36.8%&#xA;&#xA;But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>10 Sep 2026 09:29 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mv5qhszdrk2i</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3muwtwtlihk2b</link><description>OCR for Japanese manga, Swedish handwriting or Arabic print?&#xA;&#xA;There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.&#xA;&#xA;I’ve gathered 41 models into four collections, with short notes to help you choose:&#xA;&#xA;https://huggingface.co/collections/davanstrien/ocr-on-the-hub</description><pubDate>07 Sep 2026 15:43 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3muwtwtlihk2b</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mumpbzpotc2i</link><description>A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger.&#xA;&#xA;On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.&#xA;&#xA;https://huggingface.co/spaces/finebooks/bhl-ocr-leaderboard</description><pubDate>03 Sep 2026 14:53 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mumpbzpotc2i</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtz2hc57xk2x</link><description>Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>26 Aug 2026 19:19 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtz2hc57xk2x</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtwcbdgyy22y</link><description>Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub&#xA;&#xA;https://huggingface.co/datasets/biglam/britannica-illustrated-pages</description><pubDate>25 Aug 2026 17:01 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtwcbdgyy22y</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtj2qezwr22v</link><description>Every illustration in the British Library&#39;s 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved!&#xA;&#xA;One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.</description><pubDate>20 Aug 2026 10:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtj2qezwr22v</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtds3a6ho22j</link><description>Synthetic data at scale without owning a GPU: datatrove&#39;s new Jobs backend + Qwen3.8-27B → 35,837 length-controllable TL;DRs of Hugging Face cards, $0.43 per 1,000. Full guide: https://danielvanstrien.xyz/posts/2026/distilling-qwen38-datatrove-jobs/</description><pubDate>18 Aug 2026 08:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtds3a6ho22j</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3msq3gqmcys2l</link><description>FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment.&#xA;&#xA;Step one: work out which modern OCR models are actually good enough.</description><pubDate>10 Aug 2026 12:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3msq3gqmcys2l</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3msj2timahc2c</link><description>Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: https://huggingface.co/datasets/biglam/british-library-book-images&#xA;&#xA;You can also do semantic search against the images here: https://huggingface.co/spaces/davanstrien/bl-images-search</description><pubDate>07 Aug 2026 17:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3msj2timahc2c</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3ms6gi54ay22z</link><description>Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly.&#xA;&#xA;Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30.&#xA;&#xA;https://huggingface.co/datasets/huggingface/agent-usage</description><pubDate>03 Aug 2026 11:48 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3ms6gi54ay22z</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mrwuvnreds2q</link><description>A film catalogue tells you what a film is about, not what happens inside it.&#xA;&#xA;So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments.&#xA;&#xA;Search &#34;typing on a computer keyboard&#34;, land on the second it happens.</description><pubDate>31 Jul 2026 11:44 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mrwuvnreds2q</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mrtzmakxmk2n</link><description>New recipe: timestamped video captions on @hf.co Jobs.&#xA;&#xA;Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise &lt;start – end&gt; events.&#xA;&#xA;~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM). &#xA;&#xA;https://huggingface.co/datasets/uv-scripts/video</description><pubDate>30 Jul 2026 08:31 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mrtzmakxmk2n</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqmjaqkhj224</link><description>Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6.&#xA;&#xA;Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.&#xA;&#xA;huggingface.co/datasets/uv-scripts/ocr</description><pubDate>14 Jul 2026 15:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqmjaqkhj224</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqmjaqkhj224</link><description>Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6.&#xA;&#xA;Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.&#xA;&#xA;huggingface.co/datasets/uv-scripts/ocr</description><pubDate>14 Jul 2026 15:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqmjaqkhj224</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqa37fhjic2a</link><description>Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape:&#xA;https://huggingface.co/spaces/davanstrien/apollo-11-search</description><pubDate>09 Jul 2026 16:41 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqa37fhjic2a</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mq2zyveq4222</link><description>I ran 10 newer OCR models on @ai2.bsky.social&#39;s olmOCR-bench &#34;old scans&#34; subset. &#xA;&#xA;The ranking flips depending on what you actually want.</description><pubDate>07 Jul 2026 16:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mq2zyveq4222</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpoefnspvc2a</link><description>Coding agents are real users of the @hf.co Hub!&#xA;&#xA;They&#39;re searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces...&#xA;&#xA;Now there&#39;s public data: each agent&#39;s share of Hub traffic, updated monthly 👇</description><pubDate>02 Jul 2026 15:37 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpoefnspvc2a</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpnilb6js22f</link><description>The latest @commoncrawl.bsky.social crawl indexes way more than HTML pages.&#xA;&#xA;20.9M PDFs. Plus calendars, BibTeX, markdown...&#xA;&#xA;The whole index now lives in a @hf.co Bucket, so I pulled this with one SQL query straight over it (new S3 API). 2.1B rows, nothing downloaded, $0 to read.</description><pubDate>02 Jul 2026 07:19 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpnilb6js22f</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpiydikxdc2f</link><description>You can now use 100s of tools with @hf.co Buckets, thanks to the new S3 API! &#xA;&#xA;Usually just one or two lines to change.&#xA;&#xA;https://huggingface.co/docs/hub/storage-buckets-s3</description><pubDate>30 Jun 2026 12:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpiydikxdc2f</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpgo7cdfcc2u</link><description>Fairly benchmarking OCR models is hard! &#xA;&#xA;Ran a few newer OCR models on @ai2.bsky.social&#39;s olmOCR-bench &#34;old scans&#34; subset. &#xA;&#xA;Worth knowing: the score &#34;punishes&#34; models for extracting too much (letterheads, stamps, etc.)</description><pubDate>29 Jun 2026 14:11 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpgo7cdfcc2u</guid></item></channel></rss>