Skip to content

Commit 12d82da

Browse files
author
Documenter.jl
committed
build based on 9ae6539
1 parent 124ffdb commit 12d82da

2 files changed

Lines changed: 3 additions & 3 deletions

File tree

v0.1.0/.documenter-siteinfo.json

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1 +1 @@
1-
{"documenter":{"julia_version":"1.11.2","generation_timestamp":"2024-12-23T19:57:13","documenter_version":"1.8.0"}}
1+
{"documenter":{"julia_version":"1.11.2","generation_timestamp":"2024-12-23T20:03:54","documenter_version":"1.8.0"}}

v0.1.0/index.html

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2,5 +2,5 @@
22
<html lang="en"><head><meta charset="UTF-8"/><meta name="viewport" content="width=device-width, initial-scale=1.0"/><title>Home · CannotWaitForTheseOptimisers.jl</title><meta name="title" content="Home · CannotWaitForTheseOptimisers.jl"/><meta property="og:title" content="Home · CannotWaitForTheseOptimisers.jl"/><meta property="twitter:title" content="Home · CannotWaitForTheseOptimisers.jl"/><meta name="description" content="Documentation for CannotWaitForTheseOptimisers.jl."/><meta property="og:description" content="Documentation for CannotWaitForTheseOptimisers.jl."/><meta property="twitter:description" content="Documentation for CannotWaitForTheseOptimisers.jl."/><meta property="og:url" content="https://MurrellGroup.github.io/CannotWaitForTheseOptimisers.jl/"/><meta property="twitter:url" content="https://MurrellGroup.github.io/CannotWaitForTheseOptimisers.jl/"/><link rel="canonical" href="https://MurrellGroup.github.io/CannotWaitForTheseOptimisers.jl/"/><script data-outdated-warner src="assets/warner.js"></script><link href="https://cdnjs.cloudflare.com/ajax/libs/lato-font/3.0.0/css/lato-font.min.css" rel="stylesheet" type="text/css"/><link href="https://cdnjs.cloudflare.com/ajax/libs/juliamono/0.050/juliamono.min.css" rel="stylesheet" type="text/css"/><link href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/6.4.2/css/fontawesome.min.css" rel="stylesheet" type="text/css"/><link href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/6.4.2/css/solid.min.css" rel="stylesheet" type="text/css"/><link href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/6.4.2/css/brands.min.css" rel="stylesheet" type="text/css"/><link href="https://cdnjs.cloudflare.com/ajax/libs/KaTeX/0.16.8/katex.min.css" rel="stylesheet" type="text/css"/><script>documenterBaseURL="."</script><script src="https://cdnjs.cloudflare.com/ajax/libs/require.js/2.3.6/require.min.js" data-main="assets/documenter.js"></script><script src="search_index.js"></script><script src="siteinfo.js"></script><script src="../versions.js"></script><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/catppuccin-mocha.css" data-theme-name="catppuccin-mocha"/><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/catppuccin-macchiato.css" data-theme-name="catppuccin-macchiato"/><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/catppuccin-frappe.css" data-theme-name="catppuccin-frappe"/><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/catppuccin-latte.css" data-theme-name="catppuccin-latte"/><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/documenter-dark.css" data-theme-name="documenter-dark" data-theme-primary-dark/><link class="docs-theme-link" rel="stylesheet" type="text/css" href="assets/themes/documenter-light.css" data-theme-name="documenter-light" data-theme-primary/><script src="assets/themeswap.js"></script></head><body><div id="documenter"><nav class="docs-sidebar"><div class="docs-package-name"><span class="docs-autofit"><a href>CannotWaitForTheseOptimisers.jl</a></span></div><button class="docs-search-query input is-rounded is-small is-clickable my-2 mx-auto py-1 px-2" id="documenter-search-query">Search docs (Ctrl + /)</button><ul class="docs-menu"><li class="is-active"><a class="tocitem" href>Home</a></li></ul><div class="docs-version-selector field has-addons"><div class="control"><span class="docs-label button is-static is-size-7">Version</span></div><div class="docs-selector control is-expanded"><div class="select is-fullwidth is-size-7"><select id="documenter-version-selector"></select></div></div></div></nav><div class="docs-main"><header class="docs-navbar"><a class="docs-sidebar-button docs-navbar-link fa-solid fa-bars is-hidden-desktop" id="documenter-sidebar-button" href="#"></a><nav class="breadcrumb"><ul class="is-hidden-mobile"><li class="is-active"><a href>Home</a></li></ul><ul class="is-hidden-tablet"><li class="is-active"><a href>Home</a></li></ul></nav><div class="docs-right"><a class="docs-navbar-link" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl" title="View the repository on GitHub"><span class="docs-icon fa-brands"></span><span class="docs-label is-hidden-touch">GitHub</span></a><a class="docs-navbar-link" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/main/docs/src/index.md" title="Edit source on GitHub"><span class="docs-icon fa-solid"></span></a><a class="docs-settings-button docs-navbar-link fa-solid fa-gear" id="documenter-settings-button" href="#" title="Settings"></a><a class="docs-article-toggle-button fa-solid fa-chevron-up" id="documenter-article-toggle-button" href="javascript:;" title="Collapse all docstrings"></a></div></header><article class="content" id="documenter-page"><h1 id="CannotWaitForTheseOptimisers"><a class="docs-heading-anchor" href="#CannotWaitForTheseOptimisers">CannotWaitForTheseOptimisers</a><a id="CannotWaitForTheseOptimisers-1"></a><a class="docs-heading-anchor-permalink" href="#CannotWaitForTheseOptimisers" title="Permalink"></a></h1><p>Documentation for <a href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl">CannotWaitForTheseOptimisers</a>.</p><ul><li><a href="#CannotWaitForTheseOptimisers.Apollo"><code>CannotWaitForTheseOptimisers.Apollo</code></a></li><li><a href="#CannotWaitForTheseOptimisers.Muon"><code>CannotWaitForTheseOptimisers.Muon</code></a></li><li><a href="#CannotWaitForTheseOptimisers.NormGrowthCap"><code>CannotWaitForTheseOptimisers.NormGrowthCap</code></a></li></ul><article class="docstring"><header><a class="docstring-article-toggle-button fa-solid fa-chevron-down" href="javascript:;" title="Collapse docstring"></a><a class="docstring-binding" id="CannotWaitForTheseOptimisers.Apollo" href="#CannotWaitForTheseOptimisers.Apollo"><code>CannotWaitForTheseOptimisers.Apollo</code></a><span class="docstring-category">Type</span><span class="is-flex-grow-1 docstring-article-toggle-button" title="Collapse docstring"></span></header><section><div><pre><code class="language-julia hljs">Apollo(opt::AdamW = AdamW(), r::Function = dim -&gt; ceil(Int, sqrt(dim)); u = 100, sort_dims = true)
33
Apollo(η::Real, args...; kw...)
44
Apollo(arg, rank::Int; kw...)
5-
Apollo(η::Real, rank::Int; kw...)</code></pre><p>Apollo optimizer from Zhu et al. (https://arxiv.org/abs/2412.05270). Tracks moments in a low-rank subspace, aiming for Adam-like behavior with minimal additional memory usage. First argument can be an AdamW optimizer, or a learning rate (which will use the default AdamW optimizer with that learning rate). Second argument can be a rank, or a function to compute the rank from the second dimension (or the product of all dims &gt; 1) of the weight matrix (or tensor).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/02e1cb427814523cfcda8d99040b0e5d0de95710/src/rules.jl#L126-L135">source</a></section></article><article class="docstring"><header><a class="docstring-article-toggle-button fa-solid fa-chevron-down" href="javascript:;" title="Collapse docstring"></a><a class="docstring-binding" id="CannotWaitForTheseOptimisers.Muon" href="#CannotWaitForTheseOptimisers.Muon"><code>CannotWaitForTheseOptimisers.Muon</code></a><span class="docstring-category">Type</span><span class="is-flex-grow-1 docstring-article-toggle-button" title="Collapse docstring"></span></header><section><div><pre><code class="language-julia hljs">Muon(opt = AdamW(eta = 0.0003, beta = (0.9,0.95), lambda = 0.01), η = 0.02, μ = 0.95, λ = 0.01, fallback = Returns(false))
6-
Muon(; [opt, eta, mu, lambda, fallback])</code></pre><p>Muon - MomentUm Orthogonalized by Newton-schulz (https://github.com/KellerJordan/Muon)</p><p>Muon internally runs standard SGD-momentum, and then performs an orthogonalization post-processing step, in which each 2D parameter&#39;s update is replaced with the nearest orthogonal matrix using Newton-Schulz iteration.</p><p><strong>Parameters</strong></p><ul><li>Fallback optimizer (<code>opt</code>): Optimizer to use for 1D parameters or when the <code>fallback</code> function returns true</li><li>Learning rate (<code>η == eta</code>): Amount by which gradients are discounted before updating the weights</li><li>Momentum (<code>μ == mu</code>): Controls the acceleration of gradient descent in the prominent direction</li><li>Weight decay (<code>λ == lambda</code>): Controls the strength of <span>$L_2$</span> regularisation.</li><li>Fallback function (<code>fallback</code>): Function to control when, in addition to 1D arrays, the fallback optimizer should be used. Will be passed the parameter array and must return a boolean.</li></ul><p>Note: Works best with large batch sizes and may not be suitable for fine-tuning. In nanoGPT speedrun experiments, Muon is used for the internal layer &gt;2D weights, and AdamW is used for the 1D weights, embeddings, and heads.</p><p><code>Optimisers.adjust!(optimiser_state, η::Real)</code> will adjust the fallback optimizer&#39;s <code>eta</code> to <code>η * (opt.eta / eta)</code>, and Muon&#39;s <code>eta</code> to <code>η</code>, preserving their ratio, but <code>Optimisers.adjust!(optimiser, eta = η)</code> will only adjust Muon&#39;s learning rate (allowing you to adjust the fallback optimizer&#39;s learning rate separately).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/02e1cb427814523cfcda8d99040b0e5d0de95710/src/rules.jl#L4-L25">source</a></section></article><article class="docstring"><header><a class="docstring-article-toggle-button fa-solid fa-chevron-down" href="javascript:;" title="Collapse docstring"></a><a class="docstring-binding" id="CannotWaitForTheseOptimisers.NormGrowthCap" href="#CannotWaitForTheseOptimisers.NormGrowthCap"><code>CannotWaitForTheseOptimisers.NormGrowthCap</code></a><span class="docstring-category">Type</span><span class="is-flex-grow-1 docstring-article-toggle-button" title="Collapse docstring"></span></header><section><div><pre><code class="language-julia hljs">NormGrowthCap(τ = 1.01; ϵ = 1e-8, lb = 1e-7, throw = true, scale = true)</code></pre><p>Gradient norm growth limiter. <code>τ</code> controls the maximum that the gradient norm can grow from one step to the next, such that if <code>||dx||/||dx_prev|| &gt; τ</code> &amp; <code>||dx|| &gt; lb</code>, then <code>dx = dx * τ*||dx_prev||/(||dx||+ϵ)</code> Inspired by <a href="https://arxiv.org/abs/2410.01623">Chen et al.</a> and used with Apollo in <a href="https://arxiv.org/abs/2412.05270">Zhu et al.</a>, but with Optimisers.jl this will apply per-tensor instead of per-model. This implementation also introduces <code>lb</code> as a hard minimum on the gradient norm threshold, and never rescales grads below this, preventing a tensor from getting &quot;trapped&quot; near zero. This can be a fixed min, or scaled by the square root of the number of parameters in the tensor (with <code>scale = true</code>).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/02e1cb427814523cfcda8d99040b0e5d0de95710/src/rules.jl#L77-L86">source</a></section></article></article><nav class="docs-footer"><p class="footer-message">Powered by <a href="https://github.com/JuliaDocs/Documenter.jl">Documenter.jl</a> and the <a href="https://julialang.org/">Julia Programming Language</a>.</p></nav></div><div class="modal" id="documenter-settings"><div class="modal-background"></div><div class="modal-card"><header class="modal-card-head"><p class="modal-card-title">Settings</p><button class="delete"></button></header><section class="modal-card-body"><p><label class="label">Theme</label><div class="select"><select id="documenter-themepicker"><option value="auto">Automatic (OS)</option><option value="documenter-light">documenter-light</option><option value="documenter-dark">documenter-dark</option><option value="catppuccin-latte">catppuccin-latte</option><option value="catppuccin-frappe">catppuccin-frappe</option><option value="catppuccin-macchiato">catppuccin-macchiato</option><option value="catppuccin-mocha">catppuccin-mocha</option></select></div></p><hr/><p>This document was generated with <a href="https://github.com/JuliaDocs/Documenter.jl">Documenter.jl</a> version 1.8.0 on <span class="colophon-date" title="Monday 23 December 2024 19:57">Monday 23 December 2024</span>. Using Julia version 1.11.2.</p></section><footer class="modal-card-foot"></footer></div></div></div></body></html>
5+
Apollo(η::Real, rank::Int; kw...)</code></pre><p>Apollo optimizer from Zhu et al. (https://arxiv.org/abs/2412.05270). Tracks moments in a low-rank subspace, aiming for Adam-like behavior with minimal additional memory usage. First argument can be an AdamW optimizer, or a learning rate (which will use the default AdamW optimizer with that learning rate). Second argument can be a rank, or a function to compute the rank from the second dimension (or the product of all dims &gt; 1) of the weight matrix (or tensor).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/9ae6539a953f992fbe817942d4891742d8191527/src/rules.jl#L126-L135">source</a></section></article><article class="docstring"><header><a class="docstring-article-toggle-button fa-solid fa-chevron-down" href="javascript:;" title="Collapse docstring"></a><a class="docstring-binding" id="CannotWaitForTheseOptimisers.Muon" href="#CannotWaitForTheseOptimisers.Muon"><code>CannotWaitForTheseOptimisers.Muon</code></a><span class="docstring-category">Type</span><span class="is-flex-grow-1 docstring-article-toggle-button" title="Collapse docstring"></span></header><section><div><pre><code class="language-julia hljs">Muon(opt = AdamW(eta = 0.0003, beta = (0.9,0.95), lambda = 0.01), η = 0.02, μ = 0.95, λ = 0.01, fallback = Returns(false))
6+
Muon(; [opt, eta, mu, lambda, fallback])</code></pre><p>Muon - MomentUm Orthogonalized by Newton-schulz (https://github.com/KellerJordan/Muon)</p><p>Muon internally runs standard SGD-momentum, and then performs an orthogonalization post-processing step, in which each 2D parameter&#39;s update is replaced with the nearest orthogonal matrix using Newton-Schulz iteration.</p><p><strong>Parameters</strong></p><ul><li>Fallback optimizer (<code>opt</code>): Optimizer to use for 1D parameters or when the <code>fallback</code> function returns true</li><li>Learning rate (<code>η == eta</code>): Amount by which gradients are discounted before updating the weights</li><li>Momentum (<code>μ == mu</code>): Controls the acceleration of gradient descent in the prominent direction</li><li>Weight decay (<code>λ == lambda</code>): Controls the strength of <span>$L_2$</span> regularisation.</li><li>Fallback function (<code>fallback</code>): Function to control when, in addition to 1D arrays, the fallback optimizer should be used. Will be passed the parameter array and must return a boolean.</li></ul><p>Note: Works best with large batch sizes and may not be suitable for fine-tuning. In nanoGPT speedrun experiments, Muon is used for the internal layer &gt;2D weights, and AdamW is used for the 1D weights, embeddings, and heads.</p><p><code>Optimisers.adjust!(optimiser_state, η::Real)</code> will adjust the fallback optimizer&#39;s <code>eta</code> to <code>η * (opt.eta / eta)</code>, and Muon&#39;s <code>eta</code> to <code>η</code>, preserving their ratio, but <code>Optimisers.adjust!(optimiser, eta = η)</code> will only adjust Muon&#39;s learning rate (allowing you to adjust the fallback optimizer&#39;s learning rate separately).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/9ae6539a953f992fbe817942d4891742d8191527/src/rules.jl#L4-L25">source</a></section></article><article class="docstring"><header><a class="docstring-article-toggle-button fa-solid fa-chevron-down" href="javascript:;" title="Collapse docstring"></a><a class="docstring-binding" id="CannotWaitForTheseOptimisers.NormGrowthCap" href="#CannotWaitForTheseOptimisers.NormGrowthCap"><code>CannotWaitForTheseOptimisers.NormGrowthCap</code></a><span class="docstring-category">Type</span><span class="is-flex-grow-1 docstring-article-toggle-button" title="Collapse docstring"></span></header><section><div><pre><code class="language-julia hljs">NormGrowthCap(τ = 1.01; ϵ = 1e-8, lb = 1e-7, throw = true, scale = true)</code></pre><p>Gradient norm growth limiter. <code>τ</code> controls the maximum that the gradient norm can grow from one step to the next, such that if <code>||dx||/||dx_prev|| &gt; τ</code> &amp; <code>||dx|| &gt; lb</code>, then <code>dx = dx * τ*||dx_prev||/(||dx||+ϵ)</code> Inspired by <a href="https://arxiv.org/abs/2410.01623">Chen et al.</a> and used with Apollo in <a href="https://arxiv.org/abs/2412.05270">Zhu et al.</a>, but with Optimisers.jl this will apply per-tensor instead of per-model. This implementation also introduces <code>lb</code> as a hard minimum on the gradient norm threshold, and never rescales grads below this, preventing a tensor from getting &quot;trapped&quot; near zero. This can be a fixed min, or scaled by the square root of the number of parameters in the tensor (with <code>scale = true</code>).</p></div><a class="docs-sourcelink" target="_blank" href="https://github.com/MurrellGroup/CannotWaitForTheseOptimisers.jl/blob/9ae6539a953f992fbe817942d4891742d8191527/src/rules.jl#L77-L86">source</a></section></article></article><nav class="docs-footer"><p class="footer-message">Powered by <a href="https://github.com/JuliaDocs/Documenter.jl">Documenter.jl</a> and the <a href="https://julialang.org/">Julia Programming Language</a>.</p></nav></div><div class="modal" id="documenter-settings"><div class="modal-background"></div><div class="modal-card"><header class="modal-card-head"><p class="modal-card-title">Settings</p><button class="delete"></button></header><section class="modal-card-body"><p><label class="label">Theme</label><div class="select"><select id="documenter-themepicker"><option value="auto">Automatic (OS)</option><option value="documenter-light">documenter-light</option><option value="documenter-dark">documenter-dark</option><option value="catppuccin-latte">catppuccin-latte</option><option value="catppuccin-frappe">catppuccin-frappe</option><option value="catppuccin-macchiato">catppuccin-macchiato</option><option value="catppuccin-mocha">catppuccin-mocha</option></select></div></p><hr/><p>This document was generated with <a href="https://github.com/JuliaDocs/Documenter.jl">Documenter.jl</a> version 1.8.0 on <span class="colophon-date" title="Monday 23 December 2024 20:03">Monday 23 December 2024</span>. Using Julia version 1.11.2.</p></section><footer class="modal-card-foot"></footer></div></div></div></body></html>

0 commit comments

Comments
 (0)