Appearance
AI Studio - the problem, the design and how the pieces fit
Read this first. It explains why AI Studio exists, how it is built, and which guide to open for each job.
The problem
Teams that add AI to a business application hit the same six problems, one after another.
- Every application wires its own model. A vendor name and an API key end up in code. Switching vendor, or letting a customer use their own model, means a code change and a release.
- Answers are not grounded. A model that has not read your policies invents plausible ones. Feeding it documents needs chunking, embeddings, a vector index and ranking, which is a project of its own.
- Prompts are unmanaged. They live in source files and chat threads. Nobody can say which wording produced last week's bad answer, or roll back.
- Nobody can tell whether a change made things better. A new prompt or model "feels" better. There is no test set, no score and no comparison.
- Safety and cost are afterthoughts. Personal data goes to a third party, a user pastes a prompt-injection, a loop burns a month's budget, and there is no record of who called what.
- AI cannot be shipped like the rest of the application. A plugin can ship pages and workflows, but its prompts, knowledge and agents are set up by hand in each tenant.
AI Studio removes them with one rule: AI is configuration, not code. Each piece of AI is a small, versioned definition that a designer builds in Studio, a plugin ships, and any application calls by name.
The design
Six principles
| Principle | What it means in practice |
|---|---|
| Generic | Any model that speaks an OpenAI-style API works: OpenAI, Mistral, Hugging Face endpoints, Ollama, vLLM, or a tenant's own server. Embeddings work with the same providers, not only one local runtime. |
| Tenant-owned | A tenant registers its own providers, models and keys. Keys are stored as secrets, never returned by the API. |
| Profiles, not models | Assets name a model profile such as default. The tenant decides which real model, parameters, fallback and key sit behind it. |
| Versions and labels | Prompts and RAG pipelines have immutable versions. Labels such as prod and staging point at a version. Callers ask for a label, so a rollback is moving a label. |
| The engine orchestrates, the model decides | Our deterministic workflow and pipeline engines control the path, permissions and stop conditions. The model does one bounded thing at a time. Frameworks such as LangGraph would sit behind an adapter; none is required. |
| Measured | Every call is recorded (tokens, latency, cost). Evaluation datasets score a change before it goes to prod. |
The pieces
+-------------------- AI Studio (one workspace) --------------------+
| |
Models ---------+--> Model profiles ("default", "fast", "embed") |
(providers, | ^ ^ ^ |
keys, fallback)| | | | |
| Prompts (versions) Knowledge bases Guardrails |
| ^ (chunks + vectors) ^ |
| | ^ | |
| RAG pipelines <--------+ | |
| ^ | |
| Agents (tools, knowledge, guardrails) -+ |
| |
| Evaluations score prompts and pipelines. Usage records calls. |
+--------------------------------------------------------------------+
| | |
Workflow nodes OpenAI-compatible Plugins ship the
(ai.prompt ...) gateway + app keys definitions (metadata/)| Piece | Job | Guide |
|---|---|---|
| Models and providers | Register any model; map profiles to it; fallback; cost | Models and providers |
| Knowledge bases and embeddings | Turn documents into searchable chunks | Knowledge bases and embeddings |
| Prompts | Versioned wording with variables, labels, playground | Prompts |
| RAG pipelines | Answer a question from knowledge with citations | RAG pipelines |
| Agents | A model that may use tools and knowledge, within guardrails | Agents |
| AI workflow nodes | AI steps inside a business workflow | AI workflow nodes |
| Evaluations and guardrails | Prove a change is better; keep answers safe | Evaluations and guardrails |
| Usage, app keys, the gateway | Call from any code; watch cost | Usage, app keys and the gateway |
| Ship it in a plugin, CLI, SDK, MCP | Package and automate | Use AI Studio from your plugin |
What happens on a question
A user asks a RAG pipeline "How many leave days do I get?".
- Input guard checks the text (personal data, injection, blocked terms, length) and flags, masks or blocks it.
- Rewrite (optional) turns a chatty follow-up into a standalone search query.
- Retrieve searches the named knowledge bases: vector, keyword or hybrid (both, fused by rank).
- Rerank (optional) asks a model to reorder the best passages.
- Context packs the passages into a size budget.
- Generate fills the answer prompt and calls the model profile. If the profile fails, its fallback profile is tried.
- Output guard checks the answer. If nothing relevant was found the pipeline returns its
noAnswerTextinstead of guessing.
The result carries the answer, citations and a per-stage report, and one usage row is written per model call.
Where things live
- Each tenant's AI data is in that tenant's own database, in the tenant's common data area, never mixed with another tenant's. This covers model providers and profiles the tenant added, keys and credentials, prompts, RAG pipelines, knowledge bases with their documents and vectors, agents and their run history, evaluations and usage. AI Studio is a platform feature used by every application, so its tables are common not in an application's own schema.
- The shared platform database keeps only the platform's default model catalog and a small index from an app key's hash to its tenant. No secret and no tenant content is stored there.
- Vectors are stored in the platform's vector store. The column is untyped so 384, 768, 1024 and 1536 dimensions can coexist; each search filters on the dimension and the embedding model. The vector store has to be enabled for each tenant by an administrator.
- Plugins ship definitions as files under
spk-assembly/metadata/{prompt,rag,eval,knowledge,agent}/. - Permissions use the
agentresource: view, create, update, publish.
Limits you should know
Be honest with your customers about these.
- Streaming returns the whole answer as one chunk.
- Prompt caching is not used in the provider adapters.
- Knowledge ingestion supports pasted text, plain text files and web pages. PDF and Word files, crawling and scheduled re-sync are not built.
- There are no per-passage access rules yet; access is per knowledge base.
- Evaluation runs are synchronous; there is no publish gate that blocks a label move on a failing score.
- Agents have no immutable versions and the test chat is single-turn.
- Rate limits are per server instance.
- Per-tenant budgets and alerts are not built; usage and cost are shown, not enforced.
Who does what
| Role | Uses | Typical work |
|---|---|---|
| Designer (workspace admin, business analyst) | Studio, no code | Register a model, upload documents, write and label prompts, build a pipeline, set guardrails, run evaluations |
| Developer (plugin author) | Plugin Studio, CLI, SDK, MCP, files | Ship the definitions in a plugin, call them from workflows and services, test in CI |
| Application (any code) | An app API key | Call /api/v1/ai/openai/v1 with an OpenAI SDK |
Every guide below has a Designer path (clicks in Studio) and a Developer path (files, CLI, code) for the same result.
