<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><description>Machine Learning Librarian at @hf.co&#xA;</description><link>https://bsky.app/profile/danielvanstrien.bsky.social</link><title>@danielvanstrien.bsky.social - Daniel van Strien</title><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3muwtwtlihk2b</link><description>OCR for Japanese manga, Swedish handwriting or Arabic print?&#xA;&#xA;There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines.&#xA;&#xA;I’ve gathered 41 models into four collections, with short notes to help you choose:&#xA;&#xA;https://huggingface.co/collections/davanstrien/ocr-on-the-hub</description><pubDate>07 Sep 2026 15:43 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3muwtwtlihk2b</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mumpbzpotc2i</link><description>A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger.&#xA;&#xA;On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved.&#xA;&#xA;https://huggingface.co/spaces/finebooks/bhl-ocr-leaderboard</description><pubDate>03 Sep 2026 14:53 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mumpbzpotc2i</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtz2hc57xk2x</link><description>Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>26 Aug 2026 19:19 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtz2hc57xk2x</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtwcbdgyy22y</link><description>Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub&#xA;&#xA;https://huggingface.co/datasets/biglam/britannica-illustrated-pages</description><pubDate>25 Aug 2026 17:01 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtwcbdgyy22y</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtj2qezwr22v</link><description>Every illustration in the British Library&#39;s 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved!&#xA;&#xA;One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.</description><pubDate>20 Aug 2026 10:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtj2qezwr22v</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mtds3a6ho22j</link><description>Synthetic data at scale without owning a GPU: datatrove&#39;s new Jobs backend + Qwen3.8-27B → 35,837 length-controllable TL;DRs of Hugging Face cards, $0.43 per 1,000. Full guide: https://danielvanstrien.xyz/posts/2026/distilling-qwen38-datatrove-jobs/</description><pubDate>18 Aug 2026 08:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mtds3a6ho22j</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3msq3gqmcys2l</link><description>FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment.&#xA;&#xA;Step one: work out which modern OCR models are actually good enough.</description><pubDate>10 Aug 2026 12:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3msq3gqmcys2l</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3msj2timahc2c</link><description>Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: https://huggingface.co/datasets/biglam/british-library-book-images&#xA;&#xA;You can also do semantic search against the images here: https://huggingface.co/spaces/davanstrien/bl-images-search</description><pubDate>07 Aug 2026 17:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3msj2timahc2c</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3ms6gi54ay22z</link><description>Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly.&#xA;&#xA;Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30.&#xA;&#xA;https://huggingface.co/datasets/huggingface/agent-usage</description><pubDate>03 Aug 2026 11:48 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3ms6gi54ay22z</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mrwuvnreds2q</link><description>A film catalogue tells you what a film is about, not what happens inside it.&#xA;&#xA;So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments.&#xA;&#xA;Search &#34;typing on a computer keyboard&#34;, land on the second it happens.</description><pubDate>31 Jul 2026 11:44 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mrwuvnreds2q</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mrtzmakxmk2n</link><description>New recipe: timestamped video captions on @hf.co Jobs.&#xA;&#xA;Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise &lt;start – end&gt; events.&#xA;&#xA;~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM). &#xA;&#xA;https://huggingface.co/datasets/uv-scripts/video</description><pubDate>30 Jul 2026 08:31 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mrtzmakxmk2n</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqmjaqkhj224</link><description>Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6.&#xA;&#xA;Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.&#xA;&#xA;huggingface.co/datasets/uv-scripts/ocr</description><pubDate>14 Jul 2026 15:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqmjaqkhj224</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqmjaqkhj224</link><description>Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6.&#xA;&#xA;Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command.&#xA;&#xA;huggingface.co/datasets/uv-scripts/ocr</description><pubDate>14 Jul 2026 15:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqmjaqkhj224</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mqa37fhjic2a</link><description>Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape:&#xA;https://huggingface.co/spaces/davanstrien/apollo-11-search</description><pubDate>09 Jul 2026 16:41 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mqa37fhjic2a</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mq2zyveq4222</link><description>I ran 10 newer OCR models on @ai2.bsky.social&#39;s olmOCR-bench &#34;old scans&#34; subset. &#xA;&#xA;The ranking flips depending on what you actually want.</description><pubDate>07 Jul 2026 16:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mq2zyveq4222</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpoefnspvc2a</link><description>Coding agents are real users of the @hf.co Hub!&#xA;&#xA;They&#39;re searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces...&#xA;&#xA;Now there&#39;s public data: each agent&#39;s share of Hub traffic, updated monthly 👇</description><pubDate>02 Jul 2026 15:37 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpoefnspvc2a</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpnilb6js22f</link><description>The latest @commoncrawl.bsky.social crawl indexes way more than HTML pages.&#xA;&#xA;20.9M PDFs. Plus calendars, BibTeX, markdown...&#xA;&#xA;The whole index now lives in a @hf.co Bucket, so I pulled this with one SQL query straight over it (new S3 API). 2.1B rows, nothing downloaded, $0 to read.</description><pubDate>02 Jul 2026 07:19 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpnilb6js22f</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpiydikxdc2f</link><description>You can now use 100s of tools with @hf.co Buckets, thanks to the new S3 API! &#xA;&#xA;Usually just one or two lines to change.&#xA;&#xA;https://huggingface.co/docs/hub/storage-buckets-s3</description><pubDate>30 Jun 2026 12:18 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpiydikxdc2f</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mpgo7cdfcc2u</link><description>Fairly benchmarking OCR models is hard! &#xA;&#xA;Ran a few newer OCR models on @ai2.bsky.social&#39;s olmOCR-bench &#34;old scans&#34; subset. &#xA;&#xA;Worth knowing: the score &#34;punishes&#34; models for extracting too much (letterheads, stamps, etc.)</description><pubDate>29 Jun 2026 14:11 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mpgo7cdfcc2u</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mp2dclwuds22</link><description>If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about!&#xA;&#xA;I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.</description><pubDate>24 Jun 2026 16:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mp2dclwuds22</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mp2dclwuds22</link><description>If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about!&#xA;&#xA;I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.</description><pubDate>24 Jun 2026 16:24 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mp2dclwuds22</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mowyo46xdc2u</link><description>Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via &#xA;@europeana.bsky.social newspapers. &#xA;&#xA;Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>23 Jun 2026 08:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mowyo46xdc2u</guid></item><item><link>https://bsky.app/profile/danielvanstrien.bsky.social/post/3mov6jk3vks2w</link><description>I think VLM-based OCR might finally be close to working on historic newspapers!&#xA;&#xA;Many models I&#39;ve tried before failed i.e. hallucinations, repetition loops, context overflow.&#xA;&#xA;Surya OCR 2 (a 650M model!) does a very good job!</description><pubDate>22 Jun 2026 15:16 +0000</pubDate><guid isPermaLink="false">at://did:plc:7e5mpxuweopubhexwqg5l3ba/app.bsky.feed.post/3mov6jk3vks2w</guid></item></channel></rss>