What the avatar actually is

The picture does not quite match the character, or it changed, or two pictures of the same character look like two different characters. An avatar is produced by an image system that has no access to the conversation and no concept of the persona beyond a short written prompt. It illustrates the character. It is not a view of anything.

Three ways an avatar gets made

A fixed illustration set. A designer drew or commissioned a number of images and the app selects among them. Consistent by construction, limited to what exists, and free to serve. This is what most apps use for the main portrait.

A small library of variants. The same character in several poses, outfits or expressions, selected by context. Still fixed assets, more of them, and the switching is rule-based rather than generated.

Generated per request. An image model produces a new picture from a text prompt each time. Unlimited in principle, expensive per image, and inconsistent in a specific way described below.

Many products mix these: a fixed portrait for the profile and generation for anything requested during a conversation. Which combination you are using is usually apparent from whether pictures repeat exactly.

The image system has not read your conversation

This is the fact that explains the mismatches. When a picture is generated, what reaches the image model is a short text prompt — typically an appearance description assembled from the character’s settings, plus whatever was requested. Not the conversation. Not the memory. Not the persona description in full.

So the image is generated from a caption, and the caption is much smaller than the character. A detail that has been part of your conversation for months will not appear in a picture unless it is in the appearance description that the operator assembles, and that description is a short list of attributes rather than a history.

It also means the picture cannot depict what just happened. An image produced during a conversation was produced from a prompt about appearance, and any apparent connection to the topic came from words in that prompt rather than from the model understanding the exchange.

Why consistency is hard

Image models generate; they do not photograph. Each generation is an independent sample, so producing the same face twice is a problem that has to be actively solved rather than a default. The available approaches — fixing the random seed, reusing a reference image, or training a small adapter on a specific character — all reduce variation and none eliminates it, and each costs something in flexibility or money.

The practical result is a recognisable failure mode: a character whose face is roughly right and never quite the same, drifting between images the way replies drift between generations. Same underlying reason. Sampling produces variation, and consistency is the thing that requires work.

Images are the most expensive thing in the product

Text generation is billed in fractions of a cent. Image generation is billed per image, at a rate that is typically far higher, and it is charged whether or not you like the result. That gap is why images are metered almost everywhere they are offered, usually in a currency rather than a subscription allowance, and why regenerating a picture costs another unit. It is the plainest example of what a credit is converting.

The category this site will not cover

One thing belongs in any honest description of image features. Tools that generate a likeness of a real, identifiable person — a face swapped in, a photograph turned into something else, a companion built to resemble somebody who exists — are a different product category and a harmful one. They produce images of people who did not agree to them, they are the subject of law in a growing number of jurisdictions, and nothing about them appears on this site: not how they work, not what they are called, not where to find them. Stated here only so the omission is deliberate rather than an oversight.

THE PRODUCT — the avatar

  · A picture of the character
                    → an illustration, drawn from a fixed set
                      or generated from a short caption.

  · Generated images not matching the
    conversation
                    → the image model never received the
                      conversation. Only an appearance prompt.

  · A face that is never quite the same twice
                    → independent samples. Consistency is
                      work, not a default.

  · Images metered separately
                    → per-image cost far above per-reply
                      cost, charged on every attempt.

  · Which method, which model, what it costs per
    image
                    → THE OPERATOR DECIDES, and can switch
                      image models without notice.

  · Whether generated images survive in an export
                    → VARIES BY APP, and media is often the
                      first thing omitted.

What you can check

See whether pictures repeat exactly. If they do, you are looking at fixed assets and there is nothing to be inconsistent. If they never repeat, generation is happening and variation is inherent.

Find the appearance settings. Where an app exposes the attribute list its prompts are built from, that list is the actual determinant of what pictures look like, and editing it is the only real control available.

Check whether image credits expire and whether failed generations are charged. Both are terms questions and both are commonly answered in the wording about virtual currency. It is the same clause that governs balances on cancellation.

Ask whether images appear in an export. Media is frequently excluded or delivered as links that expire, which matters if a picture is something you would want to keep. Exports are narrower than people expect.

What this doesn’t tell you

It does not tell you which method any specific app uses or what it charges, and no named product is described.

It does not tell you how to generate anything, prompt anything, or improve any result. That is a builder’s subject and it is not covered here in any form.

And it does not tell you that the picture shows you the character. There is nothing to show. The character is a description and a conversation reassembled each turn, and the image is a separate illustration of a caption drawn from it.