Charlie O’Neill

Papers

Research papers, followed by shorter technical notes written with the team at Base Labs and Baseten.

Papers & preprints

1
Post-Training Science for Supervised Fine-Tuning
Charles O’Neill, Mudith Jayasekara and Harry Partridge · arXiv preprint, 2026
2
Can a Language Model Learn Facts Continually in Its Weights?
Charles O’Neill · arXiv preprint, 2026
3
Still: Amortized KV Cache Compaction in a Single Forward Pass
Charles O’Neill, Alex Sandomirsky, Harry Partridge, Mudith Jayasekara and Max Kirkby · arXiv preprint, 2026
4
DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning
Alexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani, Charles O’Neill, et al. · ICML, 2026
5
Resurrecting the Salmon: Rethinking Mechanistic Interpretability with Domain-Specific Sparse Autoencoders
Charles O’Neill, Mudith Jayasekara and Max Kirkby · arXiv preprint, 2025
6
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
Charles O’Neill, Slava Chalnev, Chi Chi Zhao, Max Kirkby and Mudith Jayasekara · arXiv preprint, 2025
8
From superposition to sparse codes: interpretable representations in neural networks
David Klindt, Charles O’Neill, Patrik Reizinger, Harald Maurer and Nina Miolane · arXiv preprint, 2025
9
Sparse Autoencoders for Disentangling Dense Embeddings of Scientific Concepts
Charles O’Neill · NeurIPS workshop on Foundation Models for Science (oral), 2024
10
Sparse autoencoders for dense text embeddings reveal hierarchical feature sub-structure
Charles O’Neill · NeurIPS workshop on Scientific Methods for Understanding Deep Learning, 2024
11
Steering semantic search with interpretable features from sparse autoencoders
Charles O’Neill · NeurIPS workshop on Foundation Model Interventions, 2024
12
Disentangling Dense Embeddings with Sparse Autoencoders
Charles O’Neill · arXiv preprint, 2024
14
Steering Language Generation: Harnessing Contrastive Expert Guidance and Negative Prompting for Coherent and Diverse Synthetic Data Generation
Charles O’Neill, Yuan-Sen Ting, Ioana Ciuca, Roberta Raileanu, Jack Miller and Thang Bui · arXiv preprint, 2023
15
AstroLLaMA: Towards specialised foundation models in astronomy
Tuan Dung Nguyen, Yuan-Sen Ting, Ioana Ciuca, Charles O’Neill, et al. · IJCNLP-AACL, 2023
16
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity
Jack Miller, Charles O’Neill and Thang Bui · TMLR, 2024
17
AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets
Ernest Perkowski, Rui Pan, Tuan Dung Nguyen, Yuan-Sen Ting, Sandor Kruk, Tong Zhang, Charles O’Neill, et al. · Research Notes of the AAS, 2024
18
Measuring Sharpness in Grokking
Jack Miller, Patrick Gleeson, Charles O’Neill, Thang Bui and Noam Levi · ICLR workshop: Bridging the Gap Between Practice and Theory, 2024

Technical notes · Base Labs and Baseten Research

Apr 2026
Post-training frontier legal agents with Baseten Research
with Mudith Jayasekara, Matthew Blau, Aaron Ellis-Bloor, Niko Grupen and Gabe Pereyra
Mar 2026
Towards infinite context windows: neural KV cache compaction
with Alex Sandomirsky and Harry Partridge
Mar 2026
Dense, on-policy, or both?
with Max Kirkby
Feb 2026
Distillation without the dark
with Max Kirkby and Mudith Jayasekara
Oct 2025
Lumina: building self-improving evaluation through customer-in-the-loop refinement
with Harry Partridge, Max Kirkby, Jonathon Liu, Paras Stefanopoulos and Mudith Jayasekara
Oct 2025
Attention-based attribution: what your model is actually looking at
with Jonathon Liu, Kimbrian Canavan, Max Kirkby and Mudith Jayasekara
Oct 2025
Iterative SFT: dense reward learning
with Jonathon Liu, Harry Partridge, Max Kirkby and Mudith Jayasekara
Oct 2025
Write small, learn forever: rank-1 LoRA for continual learning
with Max Kirkby, Harry Partridge and Jonathon Liu
Sep 2025
Practical LoRA research
with Max Kirkby
Feb 2025
Do transformers notice their own mistakes? Finding a linear hallucination detector inside LLMs
with Mudith Jayasekara, Max Kirkby, Sviatoslav Chalnev and Rune Chi Zhao
Jan 2025
Why mechanistic interpretability needs a paradigm inversion
with Mudith Jayasekara and Max Kirkby