Skip to content

Connect any AI model - providers, models and profiles ​

The real problem ​

Northwind Logistics (an example company) runs an HR helpdesk on your ERP. Their IT policy says employee questions may only go to a model hosted in their own cloud account, while a partner tenant is happy with OpenAI. The same helpdesk plugin must work for both, and must survive the vendor being switched next year.

If the plugin named a vendor model, each tenant would need a fork. Instead the plugin asks for the profile default, and each tenant decides what default means.

The idea in one minute ​

TermMeaningExample
ProviderAn endpoint that speaks an OpenAI-style API, plus how to authenticatenorthwind-vllm, https://llm.northwind.example/v1
ModelOne model on a provider, with its kind (chat or embedding), context size and list pricesllama-3.1-70b, chat, 128k context
ProfileA name assets use, pointing at a model with parameters and an optional fallbackdefault -> llama-3.1-70b, temperature 0.2, fallback backup

Supported provider types: OpenAI, OpenAI-compatible (Ollama, vLLM, LM Studio, LiteLLM and most gateways), Mistral, Hugging Face endpoints and self-hosted servers. Chat and embedding models use the same providers.

Designer path (Studio, no code) ​

Model catalog

Connect your own provider

  1. Open AI Studio > Models. The Catalog tab lists models the platform already offers.
  2. Open My providers > Add provider. Enter a code (northwind-vllm), a name, the base URL and, if the server needs one, the API key. The key is stored as a secret and is never shown again.
  3. Add a model to the provider: code (as the server names it), kind chat or embedding, context window, and list prices per million tokens if you want cost figures. For an embedding model also give its dimensions (for example 1024).
  4. Press Test on the model. You get a real round trip and its latency, or the exact error.
  5. Create the profile default that points at this model. Set temperature and max output. Optionally choose a fallback profile; it is used automatically when this one fails.
  6. Create an embedding profile too (for example embed) if you will build knowledge bases.

Endpoints must be https and public. Private or internal addresses are refused (this stops a tenant from pointing the server at its own network) unless the operator allows private endpoints, which you would do for a local Ollama in development.

Developer path ​

Assets never name a vendor model. In a plugin's files use the profile code:

json
{ "promptCode": "summarise-policy", "profileCode": "default", "userTemplate": "..." }

To register a provider and model from a script (all under /api/v1/agents/model-profiles, permission agent update):

bash
# provider
POST /providers            { "providerCode":"northwind-vllm","name":"Northwind vLLM","baseUrl":"https://llm.northwind.example/v1","authScheme":"Bearer" }
# model on that provider
POST /providers/northwind-vllm/models
                           { "modelCode":"llama-3.1-70b","name":"Llama 3.1 70B","kind":"chat","contextWindow":131072,"priceInPerM":0,"priceOutPerM":0 }
# profile
POST /                     { "profileCode":"default","modelId":42,"params":{"temperature":0.2},"fallbackProfileCode":"backup" }
# test
POST /models/42/test

Then check it works the way applications will use it:

bash
curl https://your-erp/api/v1/ai/openai/v1/chat/completions \
  -H "Authorization: Bearer sk-spark-..." -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Say hello"}]}'

How to verify ​

  • Test on the model shows success and latency.
  • After one call, Models > Usage shows a row with the profile, tokens and (if you entered prices) cost.
  • Stop the primary model on purpose. A call to default should still answer through the fallback, and the usage row names the profile that answered.

Common mistakes ​

  • Naming the vendor model in a prompt or agent. Use a profile code, or the plugin breaks on the next tenant.
  • A base URL that already ends with /chat/completions. Give the API root (.../v1); the path is added for you (a provider can override the chat path if the server is unusual).
  • Different embedding models in one knowledge base. Vectors from different models are not comparable. Changing the embedding model needs a re-index (see Knowledge bases).
  • Expecting costs without prices. Cost is computed from the list prices you enter; blank prices show tokens only.

Not built ​

Per-tenant budgets and alerts, and prompt caching in the provider adapters. Streaming returns one chunk.

Next ​

Knowledge bases and embeddings - give the model your documents.