<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><description>Back to France after some time in sunny California and happy Copenhagen. Mistral, Photoroom, Meta (xformers, FairScale, R&amp;D),  EyeTribe (acq) Mostly writing around AI</description><link>https://bsky.app/profile/bentheegg.bsky.social</link><title>@bentheegg.bsky.social - Benjamin Lefaudeux 🇺🇦</title><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3luvknhtolk2p</link><description>&#34;The Serial Scaling Hypothesis&#34; (https://arxiv.org/abs/2507.12549v1, Liu et al) is interesting I think, not as new as it completely looks (autoregressive models are used serially, models have depth,..) but feels like a good formalization and intuition as of where current GPT based LLMs will typically fail</description><pubDate>26 Jul 2025 21:57 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3luvknhtolk2p</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lua6f4orr22f</link><description>In the coming age of agents, I think vibe coding will die out, same lasting power as prompt engineering. For things LLMs excell at, you might as well stick to higher level directives and let it own the work, Claude Code is a good example. 1/2</description><pubDate>18 Jul 2025 09:52 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lua6f4orr22f</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3ltt7ulsrjs2j</link><description>Still not a lot of ML talk on bsky (at least in my feed), hence paper Sunday: my two most interesting recent reads&#xA;- H Nets arxiv.org/abs/2507.07955&#xA;- Energy Based Transformers arxiv.org/abs/2507.02092&#xA;https://arxiv.org/abs/2507.07955</description><pubDate>13 Jul 2025 06:14 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3ltt7ulsrjs2j</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3ltqvesbhok2m</link><description>Little bit of personal news, shared in other circles already: I&#39;m moving to Mistral in August, after three years at Photoroom. I&#39;m really proud of what we built in the ML team with relatively limited means, lasting SOTA on the existing foundations (saliency segmentation) while growing a lot on genAI</description><pubDate>12 Jul 2025 08:01 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3ltqvesbhok2m</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lsri5a7buk2n</link><description>Alex Nichol is one the rare many-hits researchers of the field, with on top of that a track record of practical models which affect the public/ship. That Meta wouldn&#39;t target him is pretty rich&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>29 Jun 2025 20:12 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lsri5a7buk2n</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lseukquinc27</link><description>Automatically generate a fused megakernel in triton.. diving in, but if it works half as well as it reads it would already be quite something. Aligns with torch.compile of course&#xA;&#xA; https://github.com/mirage-project/mirage/tree/mpk</description><pubDate>24 Jun 2025 19:49 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lseukquinc27</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lsbd675ob22d</link><description>Sharing that Photoroom open sourced _Dataroom_, as promised some time ago. &#xA;&#xA;Accompanying blog post and mini thread&#xA;https://github.com/photoroom/dataroom&#xA;https://www.photoroom.com/inside-photoroom/building-a-modern-data-stack&#xA;&#xA;1/N</description><pubDate>23 Jun 2025 10:00 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lsbd675ob22d</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lrx3eugd722i</link><description>Still haven&#39;t tried Cursor, but I recently moved from Github Copilot to Continue with Codestral (free API), and it&#39;s absurd how much better Continue with Codestral is (vs. Copilot with expensive and slow models). &#xA;&#xA;Made me realize that there is zero moat in this field, at least for Copilot.</description><pubDate>19 Jun 2025 08:14 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lrx3eugd722i</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lrkjvxgdek2a</link><description>Self adapting language models, still early but fascinating prospects. There&#39;s a dimensionality curse of course: since the dimensions the LLM can touch per generated token are very small as such, needs a massive lever / dimension reduction to be able to self improve.&#xA;&#xA;arxiv.org/pdf/2506.10943</description><pubDate>14 Jun 2025 08:29 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lrkjvxgdek2a</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lrgp5zpf5s2w</link><description>Great write up of AMD new offerings, catching the nvidia train on the software side it seems. 3x speedup on MI300X since release, was required but still great to grab&#xA;&#xA;https://morethanmoore.substack.com/p/amds-ai-future-is-rack-scale-helios?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26243d18-3049-4e78-8be0-ace5128f5e62_4000x2250.jpeg&amp;open=false</description><pubDate>12 Jun 2025 19:53 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lrgp5zpf5s2w</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lqsnvauxvs2i</link><description>datago now available with webdataset compatibility (streaming tarballs, so you get the data as it arrives). Just pip install datago and give it a whirl if you&#39;d like ? Speed without the dataloader processes, and typical ViT/DiT pre-processing baked in. &#xA;example code here https://github.com/Photoroom/datago/blob/main/python/benchmark_webdataset.py</description><pubDate>04 Jun 2025 20:37 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lqsnvauxvs2i</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lqgzgnkpcs2n</link><description>Great link with https://bsky.app/profile/dbreunig.bsky.social/post/3lqfeiqn3ks2f&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>31 May 2025 05:31 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lqgzgnkpcs2n</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lqdpnni5ls2w</link><description>SageAttention3 paper reads great, and looks like B200s just got a good value boost. QAT or PTQ-free use of FP4, I expected this to be much more complicated or come later to be honest. Only at the attention level and LLMs are most often MLP bottlenecked but stil&#xA;arxiv.org/abs/2505.11594&#xA;https://arxiv.org/abs/2505.11594</description><pubDate>29 May 2025 21:58 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lqdpnni5ls2w</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lovrfrfkz225</link><description>Striking in retrospect how some leaders position at the time of the first atomic bomb (“other countries won’t get it”), Truman for instance, then space age, then now AI age, rhyme. Same as before, I think a bunch of places are bound to be SOTA AI, ideas and progress cannot be pinned to a wall</description><pubDate>11 May 2025 15:27 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lovrfrfkz225</guid></item><item><link>https://bsky.app/profile/bentheegg.bsky.social/post/3lotwueuyrs2u</link><description>Apple readying in house server grade AI processors feels quite bizarre to me: Apple Intelligence has been underwhelming so far, so current status is probably far from end game. But lowering the code now into hardware limits future flexibility, bad timing ?&#xA;https://www.tomshardware.com/pc-components/cpus/apple-reportedly-readies-baltra-processors-for-ai-servers</description><pubDate>10 May 2025 22:00 +0000</pubDate><guid isPermaLink="false">at://did:plc:vgqqpolbhvqsasdfplrrbcsg/app.bsky.feed.post/3lotwueuyrs2u</guid></item></channel></rss>