Skip to content

Models and behavior

The character’s replies are written by an AI model, and the behavior (prompt) sets its manner: how detailed the scenes are, how it keeps the character, what to avoid. The same character sounds different depending on this pair. You can change it at any point in the chat.

  • Model — the “brain” that generates the replies. Models differ in quality, speed, reply length and price.
  • Behavior — instructions for the model, the storytelling style. For example, one behavior makes replies short and lively, another long and literary.
  • Preset — a saved “model + behavior” pair together with its generation settings.

You choose the first “model + behavior” pair on the character page, in the block. The first card is , the second is . If you have chatted with the character before, a model labeled “Last used” appears next to them — the one you chose last time. Under each model is its : click it to choose the pair. The button opens the full list of models. If you have already picked your own model, the button is called “Replace AI Model”.

The model card shows the memory size (for example, 4K or 16K) and the price of a reply: “FREE” or an amount in tokens per message. A plan badge such as Story GOD means the model is available starting from that plan.

  1. In the chat’s side panel, find the . It has at most two cards: the “model + behavior” pair that is writing the replies right now, and your saved preset marked “Preset”.

  2. Click a card — the next replies will be written by that pair.

  3. To choose a new pair, click , pick a model and click . In the second step, choose a and click . The behavior card shows its author, how many messages have been written with it, user ratings and the size of the instructions.

Your account has one saved preset. It appears in the side panel of all your chats, and you can pick it in any of them with one click. A new preset replaces the previous one. The “Create preset” button is visible while the block has fewer than two cards. If it isn’t there, choose another pair with the pencil on a card (it appears on hover) or remove the preset with the icon on its card. The trash icon only works in the chat where the preset was saved.

When choosing a model, use the All models, Free, Recommended, Tokens and My plan, the and the : “Trending”, alphabetical or by price. The shows the description, style features (for example, Creative or Immersive), model size, memory size (context) and price. If the model isn’t included in your plan, the card shows the required plan’s badge and an button. Labels show what CAICHAT recommends (“We recommend”) and what the character’s creator chose (“Author’s choice”).

  • Free models are available to everyone and don’t spend tokens.
  • Plan models are included in a subscription: the higher the plan, the more models. See Plans.
  • Token models charge tokens for every reply. The price of a reply depends on the model and on the size of the character’s profile — it is shown when you choose a model and in the character’s profile (“Message cost”).

Hover over a preset card in the side panel and click to open Model settings. The pencil next to it lets you choose another model and behavior for this card.

SettingWhat it changes
Creativity. Closer to 0 (Precise) — more accurate and predictable; closer to 2 (Creative) — bolder and more surprising.
The context: how much of the recent conversation the model reads. The maximum depends on the model and your plan, see Memory.
The maximum reply length.
The model first “thinks over” the reply. Available only for paid models and uses part of the reply limit.

The section has parameters for experienced users:

  • Top-P and Top-K — how wide a choice of words the model allows. Lower means stricter replies, higher means more varied ones.
  • Presence Penalty — a penalty for repeating topics already mentioned: the higher it is, the more readily the model brings in something new.
  • Frequency Penalty — a penalty for repeating the same words often.

Click to save the changes. Not all models support all parameters: unsupported ones are simply not shown.

Some models can “think” before replying: this improves the logic of complex scenes, but the reply arrives more slowly. Reasoning is turned on in the preset settings. For some models it is always on and can’t be turned off.

If the model has spent the whole reply limit on thinking, the chat shows “The model reached the response limit and returned only thoughts”. Increase Max Response Tokens in the settings and regenerate the reply.

In the left menu, in the Studio group, there are and Prompts sections — they collect all the models and behaviors on the platform: descriptions, the best characters on each model, creators’ recommendations. Creators also publish their own behaviors (prompts) there, and owners of their own AI models can connect them via API.

Which model should I choose?

Start with the models labeled “We recommend” or the one the character’s creator chose. If the replies seem weak, try a model one level higher: the difference in quality between models is noticeable.

How much does one message cost?

It depends on the model and on the size of the character’s profile. The price is shown when you choose a model, and in the character’s profile in the “Message cost” field. On models included in your plan, messages don’t spend tokens.

The model changed by itself

If the selected model became unavailable (for example, your subscription or tokens ran out), a hint appears in the message field. Choose another model in the chat’s side panel.

Didn’t find an answer?Write to us and we’ll help you sort it out.

This is how it looks on the site