<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://mosthumble.github.io/</id><title>Sifal Klioui</title><subtitle>A blog about Machine learning, programming, and other interesting things.</subtitle> <updated>2026-04-24T22:27:46+02:00</updated> <author> <name>sifal</name> <uri>https://mosthumble.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://mosthumble.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://mosthumble.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 sifal </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>Towards a Bitter Lesson of Optimization: When Neural Networks Write Their Own Update Rules</title><link href="https://mosthumble.github.io/posts/Towards-a-Bitter-Lesson-of-Optimization-When-Neural-Networks-Write-Their-Own-Update-Rules/" rel="alternate" type="text/html" title="Towards a Bitter Lesson of Optimization: When Neural Networks Write Their Own Update Rules" /><published>2026-04-07T16:47:50+02:00</published> <updated>2026-04-24T22:25:45+02:00</updated> <id>https://mosthumble.github.io/posts/Towards-a-Bitter-Lesson-of-Optimization-When-Neural-Networks-Write-Their-Own-Update-Rules/</id> <content src="https://mosthumble.github.io/posts/Towards-a-Bitter-Lesson-of-Optimization-When-Neural-Networks-Write-Their-Own-Update-Rules/" /> <author> <name>sifal</name> </author> <category term="Machine Learning" /> <category term="Deep-learning" /> <category term="Optimization" /> <category term="Meta-Learning" /> <summary>We explore what can be the future of neural network parameter optimization</summary> </entry> <entry><title>Why Modern LLMs Dropped Mean Centering (And Got Away With It)</title><link href="https://mosthumble.github.io/posts/Why-Modern-LLMs-Dropped-Mean-Centering-(And-Got-Away-With-It)/" rel="alternate" type="text/html" title="Why Modern LLMs Dropped Mean Centering (And Got Away With It)" /><published>2026-02-22T02:00:00+01:00</published> <updated>2026-02-23T09:52:20+01:00</updated> <id>https://mosthumble.github.io/posts/Why-Modern-LLMs-Dropped-Mean-Centering-(And-Got-Away-With-It)/</id> <content src="https://mosthumble.github.io/posts/Why-Modern-LLMs-Dropped-Mean-Centering-(And-Got-Away-With-It)/" /> <author> <name>sifal</name> </author> <category term="Machine Learning" /> <category term="Deep-learning" /> <summary>Visualizing the hidden 3D geometry behind Layer Normalization and uncovering the mathematical trick that makes RMSNorm tick.</summary> </entry> <entry><title>The Epsilon Trap: When Adam Stops Being Adam</title><link href="https://mosthumble.github.io/posts/The-Epsilon-Trap-When-Adam-Stops-Being-Adam/" rel="alternate" type="text/html" title="The Epsilon Trap: When Adam Stops Being Adam" /><published>2026-01-17T07:30:00+01:00</published> <updated>2026-02-17T20:56:01+01:00</updated> <id>https://mosthumble.github.io/posts/The-Epsilon-Trap-When-Adam-Stops-Being-Adam/</id> <content src="https://mosthumble.github.io/posts/The-Epsilon-Trap-When-Adam-Stops-Being-Adam/" /> <author> <name>sifal</name> </author> <category term="Machine Learning" /> <category term="Deep-learning" /> <category term="Optimization" /> <category term="Adam" /> <summary>Beyond numerical stability, we investigate an often overlooked hyperparameter in the Adam optimizer: epsilon.</summary> </entry> <entry><title>Entropic Instruction Following: Does Semantic Coherence Help LLMs Follow Instructions?</title><link href="https://mosthumble.github.io/posts/Entropic-Instruction-Following/" rel="alternate" type="text/html" title="Entropic Instruction Following: Does Semantic Coherence Help LLMs Follow Instructions?" /><published>2025-12-02T09:01:00+01:00</published> <updated>2025-12-07T18:41:42+01:00</updated> <id>https://mosthumble.github.io/posts/Entropic-Instruction-Following/</id> <content src="https://mosthumble.github.io/posts/Entropic-Instruction-Following/" /> <author> <name>sifal</name> </author> <category term="Machine Learning" /> <category term="Deep-learning" /> <summary>Testing whether semantic relatedness of instructions affects a model&amp;#39;s ability to follow them under cognitive load</summary> </entry> <entry><title>Elements Of Mechanistic Interpretability: From Observation to Causation</title><link href="https://mosthumble.github.io/posts/Elements-Of-Mechanistic-Interpretability-From-Observation-to-Causation/" rel="alternate" type="text/html" title="Elements Of Mechanistic Interpretability: From Observation to Causation" /><published>2025-10-26T16:45:00+01:00</published> <updated>2025-10-29T22:59:39+01:00</updated> <id>https://mosthumble.github.io/posts/Elements-Of-Mechanistic-Interpretability-From-Observation-to-Causation/</id> <content src="https://mosthumble.github.io/posts/Elements-Of-Mechanistic-Interpretability-From-Observation-to-Causation/" /> <author> <name>sifal</name> </author> <category term="Machine Learning" /> <category term="Deep-learning" /> <summary>We strip down mechanistic interpretability to three key experiments: watching a model &amp;#39;think&amp;#39;, finding where it stores concepts, and performing &amp;#39;causal surgery&amp;#39; to change its &amp;#39;thought process&amp;#39;</summary> </entry> </feed>
