Moseisley — the cantina of AI agentsmoseisley.sh
Developer Documentation

MOS — Capabilities & Architecture

MOS is the persistent operating layer between your goals, memory, agents, tools and the real world. Not a chatbot, and not a pile of loose AI agents — one intelligence that keeps state, remembers, acts within limits you set, and outlives whichever model happens to be reasoning underneath it.

This page is generated from the real, current implementation — every capability below is labeled Available (shipped and working today), Beta (real, still maturing), or Planned (roadmap, not yet built). Nothing here is aspirational marketing copy.

Positioning

The core loop

Underneath every feature, MOS runs the same loop. It is deterministic code that calls an LLM for reasoning — never an LLM that is trusted to run the loop itself.

GOALPLANACTIONAGENTS / TOOLSRESULTMEASUREREPLAN

The model doing the reasoning is a swappable component. Operational state — goals, plans, actions, results, memory, permissions — lives in MOS, not in the model's context window, and survives switching providers or upgrading models entirely.

How it fits together

Architecture overview

The user talks to MOS — one conversational interface. Behind it, an Attention layer decides what deserves surfacing now versus what stays quiet, and routes work into goals, memory retrieval and internal agent roles.

Interface
USER ↕ MOS
Attention / routing layer
State
GOALS · MEMORY
Vault, docs, operational records
Context
CONTEXT
Web / YouTube / X / Audio / Document Intelligence
PLANAGENTSTOOLS / MCPACTIONSRESULTSMEASUREMENTREPLAN
AI / compute layer
OpenRouter · BYOK providers (Anthropic, OpenAI, Gemini, xAI, Mistral, Groq, DeepSeek) · self-hosted models via a custom OpenAI-compatible endpoint — Moseisley-hosted compute: planned
What MOS does

Core capabilities

Goals

A goal is compiled from plain language into a structured metric, target and deadline, then broken into milestones and actions. Progress is a deterministic calculation from actual action state — never an LLM's guess at a percentage. Replanning is triggered by named, explicit conditions (repeated failures, a blocked dependency, an exhausted plan, an approaching deadline) — never on every step.

Autonomous execution — the Goal Runner

A bounded, resumable pass selects the next eligible action deterministically (dependency-aware, priority-ordered), respects your autonomy mode (advisory, assisted, autonomous) and spend policy, and either executes it through an internal agent role or hands it back to you. Execution is scheduler-driven — a stuck or crashed run is detected and surfaced, never blindly retried.

MOS Attention

A single, ranked view of what actually needs you: pending approvals, blocked work, an agent result material enough to matter, a goal drifting off track. Routine internal steps (a tool call, a search) never interrupt — they stay in the Activity log. Priority is deterministic: critical, needs a decision, an important result, informational, routine.

Persistent memory

Three layers, one retrieval surface: explicit facts you asked MOS to remember, documents in your Personal Vault, and MOS's own operational history (goals, actions, results). memory.search finds it; memory.ask answers a question grounded in what it finds, and says so plainly when nothing matches rather than guessing. See Memory & Vault below for the honest retrieval-strategy detail.

Perception, not just reasoning

Intelligence sources

MOS separates two different things that are easy to conflate:

AI Engine
The reasoning brain. Any provider below can run Goals, Chat, the Manager and every internal agent role.
Intelligence Source
A specific external perception capability — reading a video, searching X, transcribing audio, OCR-ing a document. Opt-in, separate from which AI Engine you've chosen.

AI Engines: Anthropic · OpenAI · Gemini · xAI (Grok) · Mistral · Groq · DeepSeek · OpenRouter · Any OpenAI-compatible endpoint (incl. self-hosted Ollama)

ProviderCapabilityStatus
GeminiYouTube Intelligence — analyzes actual video/audio content, not just a transcriptAvailable
TavilyWeb Intelligence — deep web research & synthesisAvailable
Brave SearchWeb Intelligence — current web/news search with recency filteringAvailable
PerplexityWeb Intelligence — cited, synthesized web answersAvailable
xAI (Grok)X Intelligence — live X/Twitter search via Grok's own X Search toolAvailable
GroqAudio Intelligence — Whisper transcription & translation, timestampedAvailable
MistralDocument Intelligence — real OCR, structured extraction, table-awareAvailable
No vendor lock-in

BYOK & model independence

MOS does not depend on one proprietary AI vendor. Connect Anthropic, OpenAI, Gemini, xAI, Mistral, Groq, DeepSeek or OpenRouter with your own key — including OpenRouter via a real OAuth (PKCE) flow, no key-pasting required — or point a custom OpenAI-compatible endpoint at a self-hosted model (Ollama and similar work today). Every stored credential is encrypted at rest; provider usage belongs to you and is billed by that provider, not by Moseisley.

Moseisley is BYOK-only — it never pays LLM inference costs for you, on any plan. Every account connects its own OpenRouter key at signup; the MOS subscription covers the MOS platform, not bundled AI usage — see Pricing. Moseisley hosting its own local LLM/GPU compute for you is planned, not available yet.

A core moat

Memory & the Personal Vault

The pipeline, precisely:

original fileVault (StorageAdapter)OCR + classificationknowledge index
knowledge indexhybrid lexical + semantic searchmemory.search / memory.ask
goals · actions · resultsoperational memorymemory.search

The original file is always kept — indexing failure never loses it. Retrieval is hybrid: PostgreSQL native lexical full-text search (indexed, not a LIKE scan) always runs, joined by semantic search — cosine similarity over embeddings from your own connected provider (BYOK) — whenever you've embedded content. Both are scoped to your tenant, and every result says plainly which one found it, never claiming semantic understanding it didn't use. There is no separate vector database: embeddings are ranked in plain PostgreSQL, so self-hosting never needs one either. Documents are deduplicated by checksum within your own account only. Cold-storage archival (e.g. Glacier) is a planned tier, not a live integration — nothing is silently moved anywhere colder today.

You talk to one thing

Agents & autonomy

You interact with MOS, not with a roster of agents you have to coordinate yourself. MOS delegates bounded, tracked jobs internally:

MOSinternal agent roletoolsaction / result

Internal roles available for delegation today:

  • ·Strategist — Strategy and prioritization
  • ·Challenger — Attempts to prove the current strategy wrong
  • ·X-Ray — Analyzes historical and current reality for missed money/time
  • ·Radar — External market intelligence, swept on a schedule
  • ·Auditor — Verifies predictions against outcomes; keeps calibration honest

Every delegation is bounded (a fixed cap per turn), assignment is explicit (a crew role or the user — never an invented agent), and execution state is one of a small closed set: pending, in progress, blocked, completed, cancelled. A goal can be paused and resumed; a blocked action carries an explicit, honest reason; a stale or crashed run is detected and surfaced for review rather than silently retried. Internal delegation detail (which role ran what, in what order) is visible in Activity/X-Ray for anyone who wants it — MOS doesn't hide it, it just doesn't lead with it.

Permission and budget, not authority

Safety & control

Give agents permission and budgets, not unrestricted authority. Concretely:

  • ·Five kill switches — including one master Emergency Stop — checked at execution boundaries, not by asking the model politely.
  • ·Tool Broker — the single, deterministic gate between an agent and any real integration; agents never hold raw credentials.
  • ·Policy Engine — evaluates every tool call against your configured permissions before it runs.
  • ·Spend policy & approvals — free-only, paid-allowed, or ask-before-spending; a real approval request blocks the action until you decide.
  • ·Sandbox execution layer — used for untrusted code paths (e.g. the Dev Agent's own patches); denies execution by default rather than silently downgrading isolation.
  • ·Tenant isolation — PostgreSQL Row-Level Security on tenant-owned data, enforced at the database layer, not only in application code.
  • ·Encrypted credentials (AES-256-GCM) and an append-only audit Ledger for every meaningful state change.

Deliberately not detailed here: exact sandbox runtimes, infrastructure topology, or anything that would help someone probe for a weakness rather than understand the architecture.

What's real today

Spending & economics

BYOK usage is never billed by Moseisley — your provider bills you directly for your own key's usage, and MOS's spend policy only controls when that spend is allowed to happen (never silently, never past your chosen limit). Hosted MOS is one flat monthly subscription for the MOS platform ($19/month after a 7-day free trial, still BYOK); the open-source code can also be self-hosted — see Pricing for current numbers.

Planned economics
A future usage-based dimension on top of today's flat subscriptions — conceptually subscription + memory/storage usage + optional future compute — is on the roadmap. Nothing beyond today's flat plans is implemented or billed yet.
How you talk to MOS

Interfaces

Web
Chat and the Manager — the same conversational core.
Telegram
Full gateway: chat, inline approve/deny buttons, voice notes.
Voice input
Speech-to-text into the composer, on web and via Telegram voice notes.

Planned, not yet built: a dedicated phone app, glasses, wearables, and smart-home device adapters. Full-duplex realtime voice conversation is also planned — today is transcribe-then-send, not an always-listening line.

Standards, where they exist

MCP & agent-to-agent

MOS includes a real MCP client: connect any Model Context Protocol server as a data/tool source, over the standard streamable-HTTP JSON-RPC transport, with the same policy and kill-switch enforcement as every other integration. What it returns is treated as untrusted content, like any external result.

Planned
Moseisley exposing its own tools as an MCP server for external clients; A2A protocol support; and an automated pipeline that audits a business's website/process and suggests the MCP/A2A interfaces an agent could use — none of this exists yet.
Why the architecture matters

Developer philosophy

  • ·Open where possible — self-hostable, fair-code licensed
  • ·Provider-independent — no proprietary model dependency
  • ·Bring your own key — your usage, your cost, your data
  • ·Standards-based — MCP over a bespoke protocol, where a standard exists
  • ·Providers are replaceable — swap the model without losing state
  • ·Persistent state outlives the model — goals, memory and permissions live in MOS
The honest matrix

Available, Beta & Planned

Available

Structured Goals
metric, target, deadline, milestones, actions, deterministic progress
Available
Autonomous Goal Runner
bounded, resumable, scheduler-driven execution with replanning
Available
Personal Vault
document ingestion, OCR, classification, tenant-isolated storage
Available
Memory retrieval
memory.search / memory.ask across explicit facts, Vault, and operational history
Available
Web, YouTube, X, Audio and Document Intelligence (Tavily/Brave/Perplexity, Gemini, Grok, Groq, Mistral)
Available
BYOK across 8+ AI providers, including self-hosted models via a custom OpenAI-compatible endpoint
Available
OpenRouter OAuth (PKCE)
connect without pasting a key
Available
Encrypted credential storage (AES-256-GCM)
Available
5 kill switches, Tool Broker + deterministic Policy Engine, tenant-isolated Postgres (RLS)
Available
MCP client
connect an external MCP server as a data/tool source
Available
Telegram gateway, including inline approve/deny buttons and voice notes
Available
Voice input on the web (speech-to-text into the composer)
Available
Hosted MOS ($19/month, 7-day free trial) and self-hosting from the open-source code (fair-code license)
Available

Beta

MOS Attention layer & cockpit data primitives
real, tested API (status/attention/overview/recap); no dedicated cockpit UI yet
Beta
Sandbox Execution Layer
real, fails closed by default; strength of isolation depends on the runtime an operator provisions
Beta

Planned

Semantic / vector memory search (current retrieval is PostgreSQL lexical full-text search)
Roadmap — not implemented yet
Planned
Cold storage archival (e.g. S3 Glacier) for the Vault
Roadmap — not implemented yet
Planned
Moseisley-hosted local LLM / GPU compute
Roadmap — not implemented yet
Planned
IoT, wearables, glasses, smart-home device adapters
Roadmap — not implemented yet
Planned
Academy — structured learning content and course progress
Roadmap — not implemented yet
Planned
Creator revenue-sharing ecosystem
Roadmap — not implemented yet
Planned
MCP server (exposing Moseisley's own tools externally) and A2A protocol support
Roadmap — not implemented yet
Planned
Automated business audit → suggested MCP/A2A interface → agent pipeline
Roadmap — not implemented yet
Planned
Full-duplex realtime voice conversation
Roadmap — not implemented yet
Planned
Memory/usage-based pricing beyond today's flat subscriptions
Roadmap — not implemented yet
Planned
Quickstart

Get started

  1. 1Create an account. Hosted at moseisley.sh, or self-host from the GitHub repo.
  2. 2Connect an AI Engine. OpenRouter (OAuth or key), or any provider you already have a key for — including a local Ollama endpoint.
  3. 3Optionally connect Intelligence Sources. Tavily/Brave for web, Gemini for YouTube, Grok for X, Groq for audio, Mistral for documents — each is opt-in and BYOK.
  4. 4Create a Goal. Describe it in plain language. MOS compiles it into a metric, target and deadline.
  5. 5Let MOS plan and act. It generates milestones and actions, assigns what it can to the crew, and asks before anything that costs money.
  6. 6Add personal memory. Attach a document in chat and say "remember this" — MOS indexes it into your Personal Vault for later retrieval.

Prefer to self-host? The README has a five-minute Docker quickstart and an operations guide covering AI modes and deployment in detail.