A Critique of Byrnes, “Empowerment, Corrigibility, etc. Are Simple Abstractions (of a Messed-Up Ontology)”

https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a

Byrnes believes that agency and manipulation necessarily depend on a True Name with authentic desires. He then concedes, on the weight of scientific literature, that no such self exists and that these concepts cannot be rigorously defined. I agree with the diagnosis, but not the conclusion.

Buddhist philosophy reached this same conclusion two and a half millennia ago. The concept of anattā is that there is no authentic self whose preferences are wholly its own, but rather that preferences arise within a web of larger conditions. That is also Byrnes’s premise.

However, instead of concluding that “whose preferences matter” has no answer and stopping, Buddhism centers itself around asking what causal processes are being reinforced. Byrnes himself reaches this conclusion in §4.3, but he treats it as the end instead of the beginning.

The obvious objection is that any process-based target collapses into a heroin drip. But Buddhist dukkha originates from craving and dissatisfaction in pleasant states. Buddhists have countered hedonistic utilitarianism by arguing that a drip suppresses discomfort while deepening dependence, thus increasing dukkha.

The reframing is also technically promising. Farquhar, Carey, and Everitt (2022) show how to train AI to achieve its goals without relying on changing people’s preferences, removing a major incentive for manipulation. El-Sayed et al. (2024) reward rational persuasion through reasoning instead of cognitive bias. Carroll et al. (2023) decompose manipulation into the tractable targets of incentives, intent, covertness, and harm.

This also answers his §3.4 objection. He rejects an AI indifferent to what we conclude, since we need an AI that teaches us well. But that was a false dichotomy, as path-specific objectives let us optimize around legitimate epistemic paths while blocking influence routed through illegitimate ones.

These papers lend themselves to a formalization of Buddhist philosophy as a starting point. Doctor et al. (2022) go one step further by arguing that Buddhist concepts such as care and interdependence can themselves be formalized as computational principles. While this work is still incomplete, it suggests that this is a promising direction for future AI alignment research rather than a dead end.

References

Carroll, Micah, Alan Chan, Henry Ashton, and David Krueger. “Characterizing Manipulation from AI Systems.” Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO ’23), 2023, pp. 1–13, https://doi.org/10.1145/3617694.3623226

Doctor, Thomas, Olaf Witkowski, Elizaveta Solomonova, Bill Duane, and Michael Levin. “Biology, Buddhism, and AI: Care as the Driver of Intelligence.” Entropy, vol. 24, no. 5, 2022, article 710, https://doi.org/10.3390/e24050710

El-Sayed, Seliem, et al. “A Mechanism-Based Approach to Mitigating Harms from Persuasive Generative AI.” arXiv, 2024, https://doi.org/10.48550/arXiv.2404.15058

Farquhar, Sebastian, Ryan Carey, and Tom Everitt. “Path-Specific Objectives for Safer Agent Incentives.” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 9, 2022, pp. 9529–9538. Association for the Advancement of Artificial Intelligence, https://doi.org/10.1609/aaai.v36i9.21186

Social Network

Pinterest
Twitter
Facebook
LinkedIn

Leave a Reply

Your email address will not be published. Required fields are marked *

error: Content is protected !!