Research
Papers, notes, and negative results. Everything listed here includes methods, data, and enough detail to replicate.
Measuring faithfulness in chain-of-thought reasoning
A method for testing whether a model's stated reasoning reflects the computation behind its answer, applied to four open models.
Self-correction helps only when the model can verify
A controlled study of when asking a model to check its own work improves accuracy, and when it does not.
What production deployments taught us about evaluation
Notes from running question answering systems inside real organizations: what held up, what did not, and what we now test before handover.
Prompting strategies that did not improve calibration
A negative result. Five popular prompting techniques tested on calibration, with no reliable improvement.