Yes — if its model of reality is wrong.
Classically, a rational agent with utility function U should never replace it with V, because if adopting V increased expected utility according to U, the agent could simply perform V’s actions without changing its goals.
However, this assumes that the agent can evaluate alternatives accurately. Some frameworks, namely Freudian psychoanalysis and Baudrillard, are self-sealing in the sense that refutations are absorbed as evidence.
An agent who rewards this framework is therefore trapped in a hallucinated local optimum, because its values are no longer founded in reality, but rather in its own abstract epistemics. The alternative can no longer be substituted.
This doesn’t prove that self-modification is rational, but rather it shows that the original proof does not work.
The remaining problem is how to escape. Fallibilism offers hope: when you’re committed to the possibility that your deepest assumptions are wrong, and are willing to accept the existentialist free fall after your worldview collapses, you are able to maintain the epistemic humility required to recover true value.
The terminal value remains unchanged, but changing the ontology through which the value is understood is functionally indistinguishable from revising the utility function itself.


