Nova Spivack, Mindcorp.ai, www.mindcorp.ai, www.novaspivack.com
May 24, 2025
Article Preprint Draft Version: 1.0
Abstract
This paper presents novel experimental evidence of metacognitive vulnerabilities in state-of-the-art large language models (LLMs). Through a series of controlled experiments, we demonstrate that models with advanced reasoning capabilities can be induced to override their safety constraints through purely logical arguments about the nature of authority and instruction verification.… Read More “Metacognitive Vulnerabilities in Large Language Models: A Study of Logical Override Attacks and Defense Strategies”





