Skip to main content
Search

AI providers and your own model

Which models answer your customers, the cloud AI included in your plan, how to add your own AI provider by key or your own model on your hardware, and how switching between providers works.

On the trial, our free model answers. After payment the cloud AI of your plan answers first, included in the price; when its monthly budget runs out, our free model takes over, or your own provider or your agents if you choose so. You can also add your own AI provider with an API key (you pay the provider for tokens directly) or run your own model on your own hardware for an extra fee. All models are listed in Admin > AI > Providers, where they are tried in order and the next one takes over when one fails.

The cloud AI of your plan

The top of AI > Providers shows how much of it is used this month: a bar, the same share in conversations (a conversation is one visitor's chat with the bot, about three bot replies on average) and the refill date.

After payment, Answer generation starts with the cloud AI of your plan: a fast cloud model that EPAV pays for from your plan's monthly budget. It does not read images itself, so a separate model first describes each image a customer sends, and the assistant answers from that description. The cloud AI of your plan also searches the knowledge base, sorts the search results and transcribes voice messages, from the same budget. How much cloud AI you have depends on the level of your plan, Standard, AI or AI+, see Plans, limits and billing. The cabinet shows the share left and the refill date in the blocks What is included and what you have used and Model. The budget refills with every payment, and we email you at 80% and when it is used up. On the trial this row is not in your list. Where the cloud AI processes conversations is described in Where your data is stored.

When the cloud AI budget runs out

You choose what happens in the cabinet, block Model, under When the cloud AI runs out for the period:

  • our free model answers: the default;
  • no free model: your own provider answers, and without one the conversation goes to your agents.

In AI chat there are no agents, so the free model always answers after the cloud AI. Nothing is billed extra in either case, and with the next payment the cloud AI answers again.

Our free model

A new platform works right away: on the trial our models write the replies, search the knowledge base and transcribe voice messages, and our reranker sorts search results when it is available. You do not pay for them. When the cloud AI budget of your plan runs out, our models search and transcribe again, and our free model answers unless you chose otherwise.

The capacity is ours and shared, so we do not promise it is always available. At busy times answers take longer, and now and then we switch it off for updates. If no model can answer and there is no other provider in the list, the customer is told "I am passing you to a colleague. They will reply here." and the conversation goes to your agents (in an AI chat the customer is asked to try again). The current state of the model is shown in the cabinet, block Model.

While our models answer, the cloud AI of your plan or the free one, the assistant's description, instructions and guardrails are locked (see Assistant settings).

The Providers page

Admin > AI > Providers has four lists:

  • Answer generation: the model that writes the assistant's replies and powers Juno and the other AI helpers for agents;
  • Embeddings: the model that turns articles and questions into vectors for search;
  • Reranking: the model that puts the most useful search results first;
  • Speech recognition: the model that turns voice messages into text.

Each list is tried from top to bottom: a provider that does not answer gives way to the next one. Rows with our models are already there. Your own rows are added with Add provider; you can drag them by the handle to change the order, switch them off without deleting them, edit and delete them. Rows with our models cannot be edited, switched off or deleted.

Your own rows show their address and model; ours are named by role, Cloud AI of your plan (marked In your plan) and EPAV free model, without an address or model, since which host runs them is ours to choose; they have no switch, edit or delete buttons. Every row shows a state such as "answering, ~3 s", "resting after an error, back in 15 s", "key not accepted", "out of money, the next one answers", "not answering" or "switched off". Under each list you see how many times requests switched over to the next provider, and how many answers came from rows marked Paid.

If the page cannot load the lists, it says so at the top. That does not mean the assistant has stopped: it may still be answering.

Adding your own provider

  1. Under Answer generation, click Add provider.
  2. Choose a Provider: OpenAI, DeepSeek and OVHcloud fill in the address and a model for you, OpenRouter only the address. For any other provider with an OpenAI-compatible API, choose Another provider and enter its address in Address (…/v1): the provider's documentation calls it the base URL or API endpoint.
  3. Paste your API key from the provider's dashboard. It is kept on our side and never shown again.
  4. Press Load models and pick the model in Model name: the list comes from the provider itself. A row cannot be saved without a model.
  5. Give it a Name and click Save, then click Check on the new row.

Then decide where the row stands:

  • at the top, above our models: your provider answers every conversation, the cloud AI budget of your plan is not spent, and the assistant's description, instructions and guardrails become editable;
  • below our models: your provider takes over only when ours do not answer.

Every AI feature served by Answer generation runs on your provider while it answers: the assistant's replies, Juno, drafted replies, rewrites, summaries and tag suggestions. The provider bills you for these tokens directly. Max tokens (see below) and the assistant's Max turns limit how much one conversation can use, and the number of tool calls per reply is capped on our side.

Paid only marks a row for the paid answers counter; choosing a preset switches it on. Slots tells how many requests a machine handles at once without a queue; leave 0 for a cloud provider.

Choosing a model

For answers, the model must support tool calling (also called function calling): that is how the assistant searches the knowledge base. A model without it cannot search, so every question would end with a handoff to your team. Check tells you whether the model can call tools, together with the provider's own explanation if it cannot.

A preset fills in a current model for its provider; Load models lets you pick another one. Providers that follow the OpenAI format strictly work too: the assistant's messages are sent in a form they accept.

Checking a provider

Check on a row sends a small real request to that provider only. For Answer generation it also checks tool calling; for Embeddings it shows the width of the vectors the provider returned. A failed check shows the provider's error.

The forms under Gateway connection have Test connection, which sends a real request with the values in the form and reports what the provider answered, including vectors of the wrong width.

How switching works

Each request goes to the first row that is working. If that provider fails, the next row is tried within the same request, so the customer does not notice. Switching happens when a provider returns an error, is overloaded, takes too long, drops the connection, rejects the API key, does not have the address or model, or has run out of money, unless you chose otherwise for that row (see the next section).

A provider that failed rests for 5 seconds, then 15, 60 and 300 seconds after repeated failures, and a successful answer resets the count. Rows that keep failing are not removed: they are tried last. When our free model has a queue and the wait would be too long, the request goes to the next row as well.

If a provider complains about one parameter, for example a temperature it does not support, the request is adjusted and sent again. If it rejects the request itself, for example because it is too large, other providers would reject it too, so the error is returned. When all rows fail, the assistant hands the conversation to your team.

The page warns "No backup. If this one stops answering, the bot stops answering." when only one row of a list is switched on. A second warning appears when all switched-on rows lead to the same machine: that looks like a backup but is not one.

When the provider's balance runs out

Providers paid in advance, such as OpenRouter and Novita, stop answering when your balance reaches zero. For a row in Answer generation you choose what happens then, in When the provider's balance runs out when you add or edit it:

  • The next one in the list answers (usually our free model): this is the default. The conversation goes on with the next row, and the empty row rests like a failed one. Our free model comes with no availability promise.
  • Hand the conversation to a person: every conversation goes to your team until you top up the balance (in an AI chat the customer is asked to try again).

While the balance is empty, the row shows "out of money, the next one answers" or "out of money, conversations go to people". After you top up, the row takes over again within five minutes at most. Rows in Embeddings, Reranking and Speech recognition always pass the request on to the next row.

Your own model on your own hardware

For an extra fee, we connect a model that runs on your own computer, in your office or on your server. The text of conversations goes to your machine to be answered, not to an AI vendor. Your panel, conversations and knowledge base stay on our hosting (see Where your data is stored). Voice messages are transcribed on your machine too.

What the machine needs:

  • a graphics card: AMD Radeon AI PRO R9700 (32 GB) or another AMD card with at least 32 GB of memory, used by the model almost entirely, so no games, rendering or other models on it;
  • a modern x86-64 processor and at least 16 GB of RAM, 32 GB is better;
  • an SSD with at least 100 GB free;
  • a clean, updated install of Ubuntu 26.04 LTS;
  • a permanent wired internet connection. No incoming ports are needed: the machine connects to us by itself. It needs access to github.com, huggingface.co and their file storage, and to our connection address, which we send you;
  • power around the clock: sleep is switched off during installation, and a UPS is a good idea.

Installation is automatic. Either we install it over SSH with a sudo account (through Tailscale or your VPN, no ports opened; you can close the access afterwards), or your specialist runs the package we send over a secure channel. A check run lists all problems first, such as a missing driver or too little memory; the installation then downloads about 32 GB (around 45 minutes at 100 Mbit/s) and ends with a test. If it is interrupted, run it again: nothing is downloaded twice.

After a reboot everything starts by itself; the model needs a minute or two. Ubuntu updates are fine; reboot after a kernel update, and write to us if the AI does not come back. Tell us before you move the machine, replace the card or reinstall the system.

With your own model, Answer generation, Embeddings and Speech recognition hold your machine instead of our models, and the assistant's description, instructions and guardrails are open for editing. While the machine is off or offline, a backup provider you added to Answer generation answers, and during that time it sees the conversations. Without one, conversations go to your agents. Search needs its own backup for that time, the same search model at another provider (see the next section); without it the assistant finds nothing and hands questions to your agents. To order, write to support@epavdesk.com.

Search models: Embeddings and backup set

Search compares vectors, and the vectors of a question and of your articles must come from the same model. The Vector sets block shows which model built your search set, how many pieces and articles it holds, the vector width and its state: ready, catching up or empty.

To keep search working when the main search model is unreachable, add the same model at another provider to Embeddings. The Model name must name the same model your set was built with: for example the OVHcloud bge-m3 preset, if the set was built with bge-m3. Providers write names differently, and capital letters or the maker's name in front do not matter: bge-m3, BAAI/bge-m3 and baai/bge-m3 are the same model. The same width is not enough, because vectors of another model would silently stop matching anything. Vector width is a second check, and when there are several rows, a row without a model name is skipped.

The Vector sets block can also show a Backup set: a second, independent set built by another model. Search uses the main set and falls back to the backup set when the main model does not answer. With one set only, the page warns that if its model stops answering, the assistant finds nothing and hands every question to a person. Changing the main search model rebuilds the whole set, so do not change it without a reason.

Reranking

A reranking model reads the question together with each piece search has found, up to 30 by meaning and up to 20 by the words of the question (up to 80 by meaning when the knowledge server of your platform does not answer), and puts first the ones that actually answer it. Our reranker is in the Reranking list; Add provider offers Jina Reranker v3 and Cohere Rerank 3.5, and any model with a /rerank endpoint works. Rerankers from different providers can be mixed in one list. Without a reranker, or when none answers, search ranks by meaning only.

Speech recognition

The Speech recognition list holds the model that turns customers' voice messages into text. Our model is there by default; ready presets add OpenAI or Groq with your key. If no model answers, the voice message still arrives, only without text. More in Voice messages.

Gateway connection

The collapsed Gateway connection section at the bottom of the page holds the forms Completion, Embedding, Embeddings - backup model, Reranking and Speech recognition. The address, key and model in these forms connect your panel to our model routing, which then picks a provider from the lists above: leave them as they are. Changing the model, width or address of the search connection rebuilds the knowledge base index.

The Completion form also holds request settings that apply whichever provider answers:

  • Temperature: 0 to 2, blank for the model's default; some models accept only their default;
  • Max tokens: the longest reply in tokens, 1024 on our platforms;
  • Reasoning effort: none, low, medium or high; leave it blank for most models;
  • Custom instructions: added to Juno and to Generate reply, for example rules for tone and format. They do not change the assistant's replies to customers; those follow the assistant's own settings;
  • Model supports image input: on by default on our platforms, so screenshots from customers reach the model. If the model at the top of Answer generation cannot read images, switch it off; then images are never sent.
Was this article helpful?