Skip to content
EN

Training and fine-tuning

Pre-training, post-training, LoRA and the recipes behind better models.

108 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Training and fine-tuning

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Self-Play Pretraining Without Natural Data

    The paper trains a language model on byte sequences generated by programs, while an RL-updated generator adapts to the learner’s progress. It reports transfer to unseen text, images, audio, and code, along with in-context learning.

    The approach explores whether self-generated data can teach transferable structure without natural-data gradient updates.

  2. Self-Play Pretraining with Zero Data

    The work trains a generator and learner from random initialization: the generator proposes programs for a universal Turing machine, and the learner trains on their outputs. It reports predictable reductions in zero-shot validation loss across images, text, audio, and melodies, plus in-context…

    It explores whether self-play can produce scalable pretraining without training on real data.

  3. A progression of policy-gradient methods for LLM training

    The post outlines an evolution from vanilla policy gradient and REINFORCE to PPO, GRPO, and GRPO variants, describing how REINFORCE estimates the policy gradient using sampled rollouts.

    It offers engineers a concise conceptual map of reinforcement-learning methods used to train LLMs.

  4. LoRA+ uses different learning rates for adapter matrices

    The paper analyzes how using the same learning rate for LoRA’s A and B matrices can hinder feature learning in wide models. It proposes using different learning rates to address this issue.

    It offers a practical fine-tuning change for improving feature learning in wide models.

  5. Weight Spectra Track Random-Label Memorization in LLMs

    A preliminary analysis uses WeightWatcher to monitor LLM weight spectra during training with random labels. The attention K matrix’s power-law exponent α tracks memorization across seeds, but the effect is layer-specific.

    The findings suggest a possible spectral signal for tracking when training shifts from generalization toward memorization.

  6. onPanda: Token-Level Annotation and Model Inspection

    onPanda is an open-source tool for LLM data annotation and model inspection. Its workflow supports token-level corrections, SFT and preference data, and inspection of token probabilities and decoding.

    Engineers working on alignment can explore a workflow for annotating on-policy data and debugging model outputs.

  7. TR-DPO Adds a KL Penalty to the DPO Loss

    The post describes Trust Region DPO (TR-DPO), which adds a KL-divergence penalty directly to the DPO loss to limit deviation from the reference model. The author says this stabilizes offline RL training.

    The approach may be relevant to engineers evaluating stability and distribution shift in DPO training.

  8. QLoRA enables 65B model fine-tuning on a 48GB GPU

    The QLoRA paper describes fine-tuning a 65B-parameter model on one 48GB GPU by backpropagating through a frozen 4-bit quantized model into LoRA adapters. It reports preserving full 16-bit fine-tuning task performance.

    Engineers can use this approach to reduce GPU memory requirements for fine-tuning large language models.

  9. Using GEPA to Tune a Model with Prompts and Labeled Examples

    The post describes a demo using GEPA to adapt the off-the-shelf Jev model, trained on synthetic data, to specific decision criteria by changing the prompt and labeling a few examples.

    It outlines a prompt-and-labeling approach for adapting a model to specific decision criteria.

  10. ORPO combines preference optimization with supervised fine-tuning

    The paper introduces ORPO, a reference-model-free approach that combines preference alignment with supervised fine-tuning. It uses a penalty for disfavored generations during preference-aligned SFT.

    Engineers can evaluate an alternative to separate SFT and preference-alignment stages that does not require a reference model.

  11. GuppyLM: a 9M-parameter language model

    The post points to GuppyLM, a roughly 9M-parameter language model described as talking like a small fish. The author says it can be trained from scratch in five minutes.

    The repository may offer engineers a small-scale project to explore language-model training.

  12. JEPA-Anything uses orthogonal factors for predictive modeling

    JEPA-Anything is a domain-agnostic framework that uses orthogonal predictive factorization to decompose latent targets into complementary factors. The authors say code, checkpoints, and a report are open.

    The paper describes a way to structure predictive-model capacity across different domains.

  13. TinyLoRA tests RL fine-tuning with 13 parameters

    The post describes TinyLoRA, which replaces LoRA’s low-rank matrix with a trainable vector projected through a random tensor. It reports that GRPO training with 13 parameters improved Qwen2.5-7B-Instruct scores on GSM8K, MATH500, and AIME24.

    It explores whether RL-based fine-tuning can reduce adapter parameter and memory requirements.

  14. RoLA: Low-Rank Linear Attention for Diffusion Transformers

    The paper introduces RoLA, a rotary-positioned low-rank linear attention method for Diffusion Transformers. It addresses the quadratic scaling of dense spatiotemporal self-attention and a compatibility issue between RoPE and the global branch of sparse low-rank hybrids.

    Engineers working on video generation can assess an approach to reducing an inference bottleneck in Diffusion Transformers.

  15. GAD for Black-Box On-Policy LLM Distillation

    The paper introduces Generative Adversarial Distillation (GAD), a method for on-policy, black-box distillation using only a proprietary teacher’s text outputs. A discriminator distinguishes student responses from teacher responses in a minimax game.

    It describes a distillation approach that does not require access to a teacher model’s logits or parameters.

  16. Building a Document Curation Classifier with Agents and SetFit

    The author describes using an agent, SetFit and Hugging Face Jobs to build a document-purpose classifier from 200 agent-labelled examples. It classified 191,724 FinePDFs-Edu documents for about $0.70 in inference compute.

    The workflow offers a concrete example of using a small classifier to reduce the cost of large-scale data curation.

  17. MaAI updates its English VAP model

    MaAI says its English VAP model was trained with a larger scale of training data. The linked project describes real-time software for turn-taking, backchannel, and head-nodding prediction.

    Engineers working on real-time interaction can review the project and its updated model.

  18. NoRA changes LoRA initialization by normalizing A

    The post describes Normalized Low-Rank Adaptation (NoRA), which normalizes columns of LoRA’s A matrix at initialization. It claims NoRA-init captures most of the gains without ongoing normalization and reports higher averages than several adapter methods on SFT and RLVR.

    The initialization approach may offer a practical way to improve LoRA training without continuous normalization.

  19. CERN Lecture Slides on Training LLMs

    Slides from a CERN lecture on training LLMs are available as a PDF. The post also says a recording is online.

    The slides may provide a useful overview of LLM training for engineers.

  20. Mol-JEPA Trains Molecular Models Across Multiple Modalities

    Mol-JEPA is a molecular foundation model that uses modality masking to learn from molecular structures, cellular phenotypes, binding affinities, and other data. The post reports out-of-distribution gains on small ADME datasets.

    Its multimodal training approach and reported results on small ADME datasets may inform molecular model development.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor