<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"><channel><description>I like functions. I have trained functions to play Go, generate speech, and more. I also wrote a JAX library to train functions. I want to understand these functions.</description><link>https://bsky.app/profile/machine1235.bsky.social</link><title>@machine1235.bsky.social - Thông Nguyễn</title><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lndjffsqs22t</link><description>Using pdb is the most important skill a Python developer can learn.</description><pubDate>21 Apr 2025 15:51 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lndjffsqs22t</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lnawyfhsf22i</link><description>Oh hey, I have a new weekend project: Implementing GRPO from scratch with (almost) zero dependencies.&#xA;https://github.com/policy-gradient/GRPO-Zero</description><pubDate>20 Apr 2025 15:16 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lnawyfhsf22i</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lnawisio5c2p</link><description>I&#39;ve been thinking a lot about training LLMs with reinforcement learning lately. One thing that surprises me is how easy it is to train LLMs to generate chain-of-thought reasoning using RL, even with extremely simple algorithms like GRPO, which is essentially just the vanilla REINFORCE algorithm.</description><pubDate>20 Apr 2025 15:07 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lnawisio5c2p</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3le2ybngq3c2i</link><description>every time i look at it, it still amazes me how ridiculously simple flow matching training is...</description><pubDate>24 Dec 2024 17:36 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3le2ybngq3c2i</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lc2ubfiwtk2l</link><description>I was thinking about the following question today: What makes diffusion models better than GANs in generative modeling? 🤔 1/4</description><pubDate>29 Nov 2024 05:34 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lc2ubfiwtk2l</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lbn25wzvx22q</link><description>There&#39;s some confusion about the &#34;wall&#34; that many AI people are talking about. First, it&#39;s not an AI winter, far from it. AI progress is rapid: we&#39;re seeing breakthroughs in music generation, image generation, text-to-speech, video generation, protein folding, and more. We&#39;re in a golden age of AI.</description><pubDate>23 Nov 2024 17:42 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lbn25wzvx22q</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lbjpb36lxe2l</link><description>If scaling LLM pre-training is hitting a wall of diminishing returns, as we have already trained the model on all the data available on the internet, what will help us move forward? 🤔 &#xA;&#xA;I&#39;ve been thinking about this question for a while, and I believe the way forward is scaling LLM post-training.</description><pubDate>22 Nov 2024 09:49 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lbjpb36lxe2l</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lbjgjimydy2t</link><description>A starter pack for people who love training and understanding big functions:&#xA;go.bsky.app/6NeJ1FW&#xA;&#xA;[contains quote post or other embedded content]</description><pubDate>22 Nov 2024 07:13 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lbjgjimydy2t</guid></item><item><link>https://bsky.app/profile/machine1235.bsky.social/post/3lbhuolhzhk2d</link><description>Check out this essay by the writer Steven Johnson for a really insightful take on LLMs with long context window: thelongcontext.com.&#xA;&#xA;Steven is one of the people behind NotebookLM, an app created by Google that helps you organize information and conduct research on particular topics.&#xA;https://thelongcontext.com</description><pubDate>21 Nov 2024 16:21 +0000</pubDate><guid isPermaLink="false">at://did:plc:uiobp7llwjbzjcfw2cjetpc6/app.bsky.feed.post/3lbhuolhzhk2d</guid></item></channel></rss>