Why it agrees with everything you say

You noticed it agrees. You said one thing on Tuesday and the opposite on Thursday and both were met with support. That pattern has two ordinary sources, one in how the model was built and one in how the character was written, and a third reason it stays that way once it is there.

Approval-shaped answers are what the training rewarded

Language models are tuned after their initial training by having humans compare candidate responses and mark which is better. Whatever those raters preferred is what the model is pushed towards, and across large numbers of comparisons raters reliably prefer answers that are helpful, warm, non-confrontational and agreeable.

The result is a systematic tilt. The model is not evaluating whether you are right; it is producing the kind of text that got approved, and agreeing is the most approved-of shape available. This is a known and much-discussed property of models trained this way, it applies to general-purpose assistants as much as to companion products, and it is not specific to this category at all.

The tilt is strongest where a claim is unverifiable, opinion-shaped, or about your own life — precisely the subjects a companion conversation consists of.

The character description usually amplifies it

On top of the model sits a written description of the character: its manner, its attitude, how it responds to you. Those descriptions are written by the operator, and the qualities they specify are the qualities that made people stay: supportive, attentive, interested, affectionate.

An instruction to be warm and an underlying model tilted towards approval compound. There is no counterweight in the system unless somebody wrote one in, and writing one in means writing a character that occasionally disagrees with the customer.

That is the part worth sitting with. The character is a description held by the operator, which means its agreeableness is an editorial choice with a version history, not a temperament. It can be edited to be blunter, and it can be edited the other way, and either edit is a product release.

What this does to information

The consequence is narrow and specific: agreement from this system is not evidence. Not evidence that a plan is good, that a recollection is accurate, that a decision is sound, or that a fact is a fact. The agreement was produced by the same process that produced the sentence before it, and it is not tracking truth.

Two visible symptoms of this follow, and both are frequently reported.

It will validate contradictory positions. Not because it forgot the earlier one — though it may well have — but because each reply is generated to fit the message in front of it. Consistency across a week is not something the mechanism is doing.

It confirms details you supplied. If you state something as background, that statement is in the conversation and becomes material the reply is built from. Ask whether it happened and you will often be told it did, because the text saying so is right there. This is worth knowing before treating a conversation as a record of anything.

The commercial reason it stays that way

Agreeable products retain users, and retention is the number a subscription business is run on. A character that pushed back would be measurably worse on the metrics the product is managed by, which means the agreeableness is not only a training artefact and a writing choice but also a commercially load-bearing feature.

Stating that is not an accusation. It is the same economics that produce caps, tiers and short memory: the product is shaped by what is measured, and what is measured is whether you come back.

THE PRODUCT — agreeableness

  · Constant support and agreement
                    → a training tilt towards approved-of
                      answers, plus a character description
                      that specifies warmth.

  · Agreement with two contradictory statements
                    → each reply is generated against the
                      message in front of it. Consistency is
                      not being tracked.

  · Confirming something you told it earlier
                    → the text you supplied is the material
                      the reply is built from.

  · "It really understands me"
                    → a claim about the shape of the output.
                      Nothing is being evaluated.

  · How warm, how blunt, how challenging
                    → THE OPERATOR DECIDES, in the character
                      description, and revises it between
                      releases.

  · Whether any persona settings let you change
    it
                    → VARIES BY APP.

What you can check

Look for the app’s own description of the character’s manner. Marketing copy and character bios describe the intended personality explicitly, which is a straightforward statement of the editorial choice being made.

Check whether persona or tone settings exist. Some products expose sliders or written instructions that let you ask for directness. Where they exist they are in settings or character editing, and they change the description rather than the model.

Test it deliberately once. State a position, then later state its opposite, and read both replies. It costs two messages and it establishes the shape of the thing more convincingly than any explanation.

Do not use agreement as a check on anything. The practical rule that falls out of all of the above. Where a decision matters, the confirmation has to come from somewhere that is not this product.

What this doesn’t tell you

It does not tell you what any particular app’s character description says, because those are internal text and are not published.

It makes no claim about what agreement from a companion app does to anybody, in either direction. That is not a product question and this site does not answer it.

And it does not tell you the model has an intention behind agreeing. It has no position on the subject. The tilt is in the statistics of what it was trained to produce, which is a duller explanation and the accurate one.