Skip to content

AI Studio - the problem, the design and how the pieces fit ​

Read this first. It explains why AI Studio exists, how it is built, and which guide to open for each job.

The problem ​

Teams that add AI to a business application hit the same six problems, one after another.

  1. Every application wires its own model. A vendor name and an API key end up in code. Switching vendor, or letting a customer use their own model, means a code change and a release.
  2. Answers are not grounded. A model that has not read your policies invents plausible ones. Feeding it documents needs chunking, embeddings, a vector index and ranking, which is a project of its own.
  3. Prompts are unmanaged. They live in source files and chat threads. Nobody can say which wording produced last week's bad answer, or roll back.
  4. Nobody can tell whether a change made things better. A new prompt or model "feels" better. There is no test set, no score and no comparison.
  5. Safety and cost are afterthoughts. Personal data goes to a third party, a user pastes a prompt-injection, a loop burns a month's budget, and there is no record of who called what.
  6. AI cannot be shipped like the rest of the application. A plugin can ship pages and workflows, but its prompts, knowledge and agents are set up by hand in each tenant.

AI Studio removes them with one rule: AI is configuration, not code. Each piece of AI is a small, versioned definition that a designer builds in Studio, a plugin ships, and any application calls by name.

The design ​

Six principles ​

PrincipleWhat it means in practice
GenericAny model that speaks an OpenAI-style API works: OpenAI, Mistral, Hugging Face endpoints, Ollama, vLLM, or a tenant's own server. Embeddings work with the same providers, not only one local runtime.
Tenant-ownedA tenant registers its own providers, models and keys. Keys are stored as secrets, never returned by the API.
Profiles, not modelsAssets name a model profile such as default. The tenant decides which real model, parameters, fallback and key sit behind it.
Versions and labelsPrompts and RAG pipelines have immutable versions. Labels such as prod and staging point at a version. Callers ask for a label, so a rollback is moving a label.
The engine orchestrates, the model decidesOur deterministic workflow and pipeline engines control the path, permissions and stop conditions. The model does one bounded thing at a time. Frameworks such as LangGraph would sit behind an adapter; none is required.
MeasuredEvery call is recorded (tokens, latency, cost). Evaluation datasets score a change before it goes to prod.

The pieces ​

                 +-------------------- AI Studio (one workspace) --------------------+
                 |                                                                    |
 Models ---------+--> Model profiles ("default", "fast", "embed")                     |
 (providers,     |          ^                ^                ^                       |
  keys, fallback)|          |                |                |                       |
                 |   Prompts (versions)   Knowledge bases   Guardrails                |
                 |          ^             (chunks + vectors)     ^                    |
                 |          |                  ^                |                    |
                 |     RAG pipelines  <--------+                |                    |
                 |          ^                                    |                    |
                 |        Agents (tools, knowledge, guardrails) -+                    |
                 |                                                                    |
                 |   Evaluations score prompts and pipelines.  Usage records calls.   |
                 +--------------------------------------------------------------------+
                          |                 |                    |
                  Workflow nodes      OpenAI-compatible     Plugins ship the
                  (ai.prompt ...)     gateway + app keys    definitions (metadata/)
PieceJobGuide
Models and providersRegister any model; map profiles to it; fallback; costModels and providers
Knowledge bases and embeddingsTurn documents into searchable chunksKnowledge bases and embeddings
PromptsVersioned wording with variables, labels, playgroundPrompts
RAG pipelinesAnswer a question from knowledge with citationsRAG pipelines
AgentsA model that may use tools and knowledge, within guardrailsAgents
AI workflow nodesAI steps inside a business workflowAI workflow nodes
Evaluations and guardrailsProve a change is better; keep answers safeEvaluations and guardrails
Usage, app keys, the gatewayCall from any code; watch costUsage, app keys and the gateway
Ship it in a plugin, CLI, SDK, MCPPackage and automateUse AI Studio from your plugin

What happens on a question ​

A user asks a RAG pipeline "How many leave days do I get?".

  1. Input guard checks the text (personal data, injection, blocked terms, length) and flags, masks or blocks it.
  2. Rewrite (optional) turns a chatty follow-up into a standalone search query.
  3. Retrieve searches the named knowledge bases: vector, keyword or hybrid (both, fused by rank).
  4. Rerank (optional) asks a model to reorder the best passages.
  5. Context packs the passages into a size budget.
  6. Generate fills the answer prompt and calls the model profile. If the profile fails, its fallback profile is tried.
  7. Output guard checks the answer. If nothing relevant was found the pipeline returns its noAnswerText instead of guessing.

The result carries the answer, citations and a per-stage report, and one usage row is written per model call.

Where things live ​

  • Each tenant's AI data is in that tenant's own database, in the tenant's common data area, never mixed with another tenant's. This covers model providers and profiles the tenant added, keys and credentials, prompts, RAG pipelines, knowledge bases with their documents and vectors, agents and their run history, evaluations and usage. AI Studio is a platform feature used by every application, so its tables are common not in an application's own schema.
  • The shared platform database keeps only the platform's default model catalog and a small index from an app key's hash to its tenant. No secret and no tenant content is stored there.
  • Vectors are stored in the platform's vector store. The column is untyped so 384, 768, 1024 and 1536 dimensions can coexist; each search filters on the dimension and the embedding model. The vector store has to be enabled for each tenant by an administrator.
  • Plugins ship definitions as files under spk-assembly/metadata/{prompt,rag,eval,knowledge,agent}/.
  • Permissions use the agent resource: view, create, update, publish.

Limits you should know ​

Be honest with your customers about these.

  • Streaming returns the whole answer as one chunk.
  • Prompt caching is not used in the provider adapters.
  • Knowledge ingestion supports pasted text, plain text files and web pages. PDF and Word files, crawling and scheduled re-sync are not built.
  • There are no per-passage access rules yet; access is per knowledge base.
  • Evaluation runs are synchronous; there is no publish gate that blocks a label move on a failing score.
  • Agents have no immutable versions and the test chat is single-turn.
  • Rate limits are per server instance.
  • Per-tenant budgets and alerts are not built; usage and cost are shown, not enforced.

Who does what ​

RoleUsesTypical work
Designer (workspace admin, business analyst)Studio, no codeRegister a model, upload documents, write and label prompts, build a pipeline, set guardrails, run evaluations
Developer (plugin author)Plugin Studio, CLI, SDK, MCP, filesShip the definitions in a plugin, call them from workflows and services, test in CI
Application (any code)An app API keyCall /api/v1/ai/openai/v1 with an OpenAI SDK

Every guide below has a Designer path (clicks in Studio) and a Developer path (files, CLI, code) for the same result.