Author
Shashank Kumar
0 works0 citations
Recent research
- AI & ComputingOpen access
The Provably Aligned AI Core: A Deterministic Safety Wrapper for Self-Improving Agents
Self-improving artificial intelligence—systems that modify their own source code—promises exponential capability gains but also poses an existential risk: a singleunchecked code modification could permanently remove all safety constraints.Existing alignment methods (RLHF, Constit...