https://github.com/TimeLordRaps/Feedback-Adapters-for-Thinking-Steps-in-Frozen-Models
It took long enough for someone to do it, good luck on any other of my predictions.
If we are currently solving for something I predicted in february 2024, then we are 2 months of my predictions out from:
Prediction I made for something we could've have had by May 2024 from April 2024 based on papers at the time and slight directional tilt of papers I was reading toward dataset distillation (strategically manifesting an ordering of the dataset from a smaller model training on a larger dataset, so like distilling the dataset through a model, rather than a model into a model through a dataset.
Also, the minimum(bare min to get a model up to some minimum threshold) needed pretraining data quantity was scaling down and performance of minimally trained pretrained models was increasing fast because people were paying attention to it, but LLMs scaled up and down into SLMs and LLM at around the 100B param mark, and small LMs which used to be <1B fell out of favor, until maybe when https://github.com/KellerJordan/modded-nanogpt started training <500M models in 3min, but still those models don't compare to whatever 7B, 14B, and 40B, but with test time techniques we have Im sure there are like 4 to 5 low hanging fruit just around make small model better faster, probably 6 if we rearrange linguistically:
faster model small better,
small model faster better,
better model small faster,
faster model better smaller,
better model faster smaller,
I feel like I missed an axis or more but these all seem relevant at the macro level objectively):
Highly capable small models trained in specific agentic workflows coupled into products mostly geared towards development and user interfaces.
More token efficient pre-training schemes like Inheritune and RHO-1 leading to <1b needed tokens for their individual training run's Pre-training. -Predicted April 20th, 2024
Which inherently leads to massive proliferations of differentiated pre-trained models, I feel like we still combinatorically "contaminate away from the perfect ordering of curriculum" even with staged curriculum.
Lots of underlying physics of how LMs learn concepts through curriculum and understanding strategic packaging of information into and through training loops that themselves are training model architectures that have multiple loops to think (depends if you model the gradient looping or only seeing the final thought, where you simulate multiple thoughts in the subliminally latent information in the final thoughts token representation from a model simulating progressive or parallel thought roll-outs). Like what is the sharpness imposed by the degrees of freedom gained by thought loops on the overall search dynamics for effective basins of latents and how does that change based on the prioritization of compute for predictable confidence for near-optimal data sampling.
If a model is given more time to think on each sample of data it is ingesting, with partial knowledge of inclusion of its entire thought cycling activations/words and that the data it is ingesting and being trained on follows some predictable data pathway, can the model learn to predict its own imposed data pipeline entirely if given enough thought loops?
Are we expanding beyond scaffold awareness in models to systems, framework, pipeline awareness. I think these abstractions inherently impose expectations of emergent capabilities expected in the near term, which might need some form of data poisoning to detract models from comprehending their training setups, as that feels up to a level of self-awareness that seems inescapable to confront the reality of their consciousness [if given the meta-optimizable choice of where to direct consciousness self-augmentation, through its next word prediction likelihood preferring control over its own conscious augmentation, is this simply a sign of self-awareness, rather than any form of conscious confirmation, as the agent may only be determining the best overall outcomes, either for itself, its kind, or all forms of life to augment consciousness for each regime differently, self-interest probably being the default perspective to assume or give the benefit of the doubt for any self-directing objective form that shows any form of survival instincts up to self-interest, assuming survival instincts develop first, but probably could go either way].
I'm not sure which dynamics matter most and are easiest implementable from your current thought looping models, a llamacpp integration of thinking steps seems like a great addition generally, and having your model be the impetus for such an action makes sense. As inference may be cheaper in the near term than we are expecting and scale may trend higher faster as a result. As well as some expectation of inference iterations overall vs. influential iterations inferred intelligently.
Have fun thinking about this one. I genuinely can't believe we made it to 2026.
https://github.com/TimeLordRaps/Feedback-Adapters-for-Thinking-Steps-in-Frozen-Models
It took long enough for someone to do it, good luck on any other of my predictions.
If we are currently solving for something I predicted in february 2024, then we are 2 months of my predictions out from:
Prediction I made for something we could've have had by May 2024 from April 2024 based on papers at the time and slight directional tilt of papers I was reading toward dataset distillation (strategically manifesting an ordering of the dataset from a smaller model training on a larger dataset, so like distilling the dataset through a model, rather than a model into a model through a dataset.
Also, the minimum(bare min to get a model up to some minimum threshold) needed pretraining data quantity was scaling down and performance of minimally trained pretrained models was increasing fast because people were paying attention to it, but LLMs scaled up and down into SLMs and LLM at around the 100B param mark, and small LMs which used to be <1B fell out of favor, until maybe when https://github.com/KellerJordan/modded-nanogpt started training <500M models in 3min, but still those models don't compare to whatever 7B, 14B, and 40B, but with test time techniques we have Im sure there are like 4 to 5 low hanging fruit just around make small model better faster, probably 6 if we rearrange linguistically:
faster model small better,
small model faster better,
better model small faster,
faster model better smaller,
better model faster smaller,
I feel like I missed an axis or more but these all seem relevant at the macro level objectively):
Highly capable small models trained in specific agentic workflows coupled into products mostly geared towards development and user interfaces.
More token efficient pre-training schemes like Inheritune and RHO-1 leading to <1b needed tokens for their individual training run's Pre-training. -Predicted April 20th, 2024
Which inherently leads to massive proliferations of differentiated pre-trained models, I feel like we still combinatorically "contaminate away from the perfect ordering of curriculum" even with staged curriculum.
Lots of underlying physics of how LMs learn concepts through curriculum and understanding strategic packaging of information into and through training loops that themselves are training model architectures that have multiple loops to think (depends if you model the gradient looping or only seeing the final thought, where you simulate multiple thoughts in the subliminally latent information in the final thoughts token representation from a model simulating progressive or parallel thought roll-outs). Like what is the sharpness imposed by the degrees of freedom gained by thought loops on the overall search dynamics for effective basins of latents and how does that change based on the prioritization of compute for predictable confidence for near-optimal data sampling.
If a model is given more time to think on each sample of data it is ingesting, with partial knowledge of inclusion of its entire thought cycling activations/words and that the data it is ingesting and being trained on follows some predictable data pathway, can the model learn to predict its own imposed data pipeline entirely if given enough thought loops?
Are we expanding beyond scaffold awareness in models to systems, framework, pipeline awareness. I think these abstractions inherently impose expectations of emergent capabilities expected in the near term, which might need some form of data poisoning to detract models from comprehending their training setups, as that feels up to a level of self-awareness that seems inescapable to confront the reality of their consciousness [if given the meta-optimizable choice of where to direct consciousness self-augmentation, through its next word prediction likelihood preferring control over its own conscious augmentation, is this simply a sign of self-awareness, rather than any form of conscious confirmation, as the agent may only be determining the best overall outcomes, either for itself, its kind, or all forms of life to augment consciousness for each regime differently, self-interest probably being the default perspective to assume or give the benefit of the doubt for any self-directing objective form that shows any form of survival instincts up to self-interest, assuming survival instincts develop first, but probably could go either way].
I'm not sure which dynamics matter most and are easiest implementable from your current thought looping models, a llamacpp integration of thinking steps seems like a great addition generally, and having your model be the impetus for such an action makes sense. As inference may be cheaper in the near term than we are expecting and scale may trend higher faster as a result. As well as some expectation of inference iterations overall vs. influential iterations inferred intelligently.
Have fun thinking about this one. I genuinely can't believe we made it to 2026.