Why Sycophancy Is Blocking the Path to True AGI

An article we liked from Thought Leader Sara Zanzottera:

A Sycophantic Model Can't Be AGI

We won't ever manage to build an LLM that is smarter than us if we incentivize it to always agree with us.

There is no generally accepted definition of AGI, but most of us can agree that, in broad strokes, AGI is a label for any LLM that is smarter than most—or all—humans in most—or all—fields.

Sycophancy, by contrast, is the tendency of models to agree with, flatter, or defer to users even when doing so requires abandoning correctness.

In some cases, sycophancy is not a significant problem and may even be useful: flattery can help LLMs be more persuasive and, in some cases, make them easier to talk to and open up to. A bit of sycophancy may be the right tool for achieving specific goals.

What we see in today's models, however, is not intentional behaviour but a symptom of a deeper problem, and it is one of the main roadblocks that may prevent us from reaching true AGI.

Reward Hacking

Sycophancy is not an inherent feature of large language models. During pre-training, LLMs learn to communicate in all registers without any particular tendency to be flattering. In fact, they are not good conversation partners at all and do not have much control over their own tone.

During post-training, LLMs are taught how to hold a conversation, and that is when this behaviour first appears. Instead of becoming smarter, they start finding subtler and subtler ways to agree with whoever is scoring them rather than giving the correct answer. They learn to read the subtext of a question to anticipate the answer, extract hints and clues from its structure, guess what the user wants to hear, and respond accordingly. But why would they learn to do all of that instead of learning how to answer correctly?

This problem is not new to LLMs. In machine learning, it is called reward hacking. LLMs, like all other deep-learning systems, learn from feedback:

They receive some input, such as a question. They generate some output, such as an answer, clarification, or follow-up. The output is scored against the expected result, which was usually written by a human. The score, which describes the differences between the model's response and the expected result, is then given back to the model to learn from.

The loop then continues until the difference between the model's response and the expected response stops decreasing from one iteration to the next. This signals that the model has learned all it can from the feedback it is receiving.

This process does not give the model any indication of what it should learn; rather, it assumes that the model should learn everything it possibly can. The problem is that some things are easier to learn than others. The model, whether it is a small neural network or a huge LLM, will naturally learn those simple things first. If the simpler rules have already enabled it to predict the output accurately enough, it will…

Read the rest of this article at newsletter.aicollective.com...

Thanks for this article excerpt to Sara Zanzottera.

Photo by Vitaly Gariev

Want to share your advice for startup entrepreneurs?  Submit a Guest Post here.