What AstroPT knows about galaxies, and what that can teach us about LLMs
Abstract
Using a galaxy-image transformer with known physical concept ordering, the study shows that linear probes recover real structure and that concepts emerge in a fixed difficulty-based sequence during training.
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.
Community
Large language models are hard to interpret in part because their training data is broad, messy, and not grounded in known relations. Scientific foundation models promise a cleaner setting. In astronomy, many relationships in the data are already known from theory, observation, and long-standing empirical work. This perspective permits us to ask two questions which we attempt to answer in this work: when such a model is trained, does its hidden space recover known scientific structure? And can we use that structure as a tool to probe the model’s learning dynamics and internalized knowledge?
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Foundation Models for Astrophysics (2026)
- A Multimodal Approach to Star--Galaxy Separation using SPHEREx Spectrophotometry and DESI Legacy Survey Imaging (2026)
- What's Missing in AGN Feedback? Lessons learnt from Magneticum, IllustrisTNG and Simba (2026)
- FLAGS II: Constraining Galaxy Formation Models with Dimensionality Reduction of Direct Observables (2026)
- Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning (2026)
- Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints (2026)
- Six-Class BPT Galaxy Classification for Survey-Scale AGN Candidate Prioritization: Deep Tabular Model and Informative Missingness Signals (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.22614 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper