Post
14
getting ready to fine tune 3.8
- evolutionary strategies
- behavior steering experiments
- expanded dataset
- bringing back ORPO
- more orthogonal evals to keep overfitting minimum
- most probably will take abliterations as base, either mine or somebody else's
- random entropy addition from huggingface fine tunes (take what is popular on hf and randomly introduce into the lineage)
- bring more vibe coding: turns out LLMs know how to fine tune
- evolutionary strategies
- behavior steering experiments
- expanded dataset
- bringing back ORPO
- more orthogonal evals to keep overfitting minimum
- most probably will take abliterations as base, either mine or somebody else's
- random entropy addition from huggingface fine tunes (take what is popular on hf and randomly introduce into the lineage)
- bring more vibe coding: turns out LLMs know how to fine tune