Does a language model respond to the person in front of it, or to the category that person falls into?
Computer science and symbolic systems at Stanford. Investor at S32. tanekim@stanford.edu
I work on alignment. The question above is the one I keep arriving at, and I have now asked it three times in three different media.
At sixteen I studied what happens to young writers psychologically when their interior work is made legible to an audience. Then I spent five years building publications that gave half a million readers access to that work. I now think the same structure is the operative variable in how AI systems fail: a model that tracks the category rather than the person will be fluent, appropriate, and wrong in a specific way — and the failure has the same shape whether the object is a task, a principal, or the world.
Before this I built low-latency systems for equities, ran AI deployment for portfolio companies, and invested in deep tech. That work taught me how to evaluate things. It is not what I am doing now.
Four experiments testing whether frontier models track individuals or categories, each with a control the field does not currently apply: token-matched irrelevant history, length-matched verbosity, an instruction-recovery arm, and a perception control on incentive elasticity. Base and instruct checkpoints compared directly.
Argues that the catalogued alignment failures — reward hacking, sycophancy, deceptive alignment, goal misgeneralization, instrumental convergence — are one failure: the system's representation of its object has its fidelity criterion set by the objective rather than by the object. Proves the resulting blindness is bounded by the objective's channel capacity and invariant to scale, that a never-violated commitment carries no information about the disposition beneath it, and that the disposition can only be selected for by a party already modeling the system that way.
Isolated the psychological effects of young writers publishing their work and being read by a broader audience.
Comparative analysis of how mediums of artistic expression have shifted across American history, and what automation does to them.
A low-dimensional harness representation for autonomous agents — latent memory and chat context in a compressed form.
Automation and thinking tools, built and open sourced.