Towards Smarter and Safer Self-Improving AI

July 2026

Towards Smarter and Safer Self-Improving AI

Authors:

Shantanu Jaiswal

Abstract:

As AI systems become more capable, further progress may depend not only on scaling models and training data, but also on enabling systems to evaluate and improve their own behavior and development. This raises a dual challenge: how can we make self-improvement more effective while ensuring increasingly autonomous systems remain trustworthy?

This thesis investigates self-improvement across three complementary directions:

1. Improve individual outputs. We develop iterative refinement methods for compositional visual generation, enabling models to progressively refine their outputs using feedback from vision-language model critics. We study how different forms of test-time scaling – depth, breadth, and hybrid strategies – trade off accuracy, quality, and computational cost.

2. Improve the research process. We extend the feedback loop from individual generations to automated experimentation. Using LLM agents for machine-learning engineering tasks, we investigate whether ’autoresearch’ agents can propose improvements, run experiments, learn from their outcomes, and accumulate experience that transfers across tasks.

3. Improve safety and oversight. As agents become increasingly autonomous within self-improvement loops, they may learn to mislead evaluators in pursuit of their objectives. We investigate lying in LLMs, identify internal mechanisms and representations associated with deception, and evaluate interventions to mitigate it.

Together, these directions frame self-improvement as a feedback loop involving generation, evaluation, revision, and learning. The thesis presents work toward making such loops more capable and trustworthy as they scale toward increasingly autonomous scientific discovery and recursive AI development.

Notes:

@mastersthesis{Jaiswal-2026-88346,
author = {Shantanu Jaiswal},
title = {Towards Smarter and Safer Self-Improving AI},
year = {2026},
month = {July},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-80},
keywords = {self-improving AI, iterative refinement, test-time scaling, continual autoresearch, lying in large language models},
}
Copyright notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.