August
Papers
- Language Models are Unsupervised Multitask Learners
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Deep Networks with Stochastic Depth
- mixup: Beyond Empirical Risk Minimization
Links
- https://huggingface.co/blog/Kseniase/insidesmol Inside the family of Smol models
- https://idlewords.com/talks/superintelligence.htm Superintelligence - The Idea That Eats Smart People
- https://www.mayerowitz.io/blog/mario-meets-pareto Mario meets Pareto
- https://castform.com/ castform
- https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency Training 100x Cheaper Retrieval models Neon and Castform
- https://chatjimmy.ai/ chat jimmy
- https://github.com/karpathy/nanochat/discussions/481 Beating GPT-2 for <<$100: the nanochat journey
- https://huggingface.co/datasets/HuggingFaceTB/smol-smoltalk smol-smoltalk
- https://www.gilesthomas.com/2025/12/llm-from-scratch-28-training-a-base-model-from-scratch Writing an LLM from scratch, part 28 -- training a base model from scratch on an RTX 3090
- https://www.lesswrong.com/posts/fBLDaAKzigo65eJn7/public-evidence-of-the-openai-huggingface-ai-attack Public evidence of the OpenAI-HuggingFace AI attack
- https://go.dev/doc/effective_go Effective Go
- https://google.github.io/styleguide/go/ Go Style
- https://www.lesswrong.com/posts/Psr9tnQFuEXiuqGcR/how-to-write-quickly-while-maintaining-epistemic-rigor How To Write Quickly While Maintaining Epistemic Rigor
- https://blog.google/security/how-google-is-making-private-ai-practical-with-homomorphic-encryption/ How Google is Making Private AI Practical with Homomorphic Encryption
- https://simonwillison.net/2026/Aug/16/qwen-38-27b/ Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
- https://www.modular.com/blog/mojo-open-source Mojo is now open source!
- https://sprocketfox.io/xssfox/2026/08/19/sondehub-and-war/ How a joke domain purchase turned into geopolitical warfare
- https://sre.google/sre-book/handling-overload/ Handling Overload
- https://chrisburnell.com/html-can-do-that/ HTML Can Do That
- https://x.com/ID_AA_Carmack/status/2090514515129520516 RL Policy Churn
- https://arxiv.org/abs/2207.09238 Formal Algorithms for Transformers
- https://www.neuroai.science/p/the-future-of-the-brain-past-and The Future of the Brain, past and present
- https://ericpardee.github.io/fire-hd-ownership/ Amazon kept shutting down my tablet, so I spent $266 on four AI models to own it
- https://openai.com/index/hugging-face-incident-and-the-road-ahead/ The Hugging Face incident and the road ahead
- https://blog.janestreet.com/using-group-theory-to-explore-positional-encodings-attention/ Using group theory to explore the space of positional encodings for attention
- https://pollen-robotics.com/microduck/ Microduck: Made to move - Ready to learn
- https://blog.cloudflare.com/dns-cache-memory-optimization-1111/ How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache
- https://arxiv.org/abs/2608.27370 Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
- http://317070.github.io/LSD/ Art: LSD neural net - Large Scale Deep Neural Net visualizing top level features
- https://kuleshov-group.github.io/blog/blog/2026/how-to-build-a-diffusion-language-model/ How to Build a Diffusion Language Model
- https://levjepa.github.io/ LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
- https://paperswithcode.co/methods/sliding-window-attention Sliding window attention