MOS — Capabilities & Architecture
MOS is the persistent operating layer between your goals, memory, agents, tools and the real world. Not a chatbot, and not a pile of loose AI agents — one intelligence that keeps state, remembers, acts within limits you set, and outlives whichever model happens to be reasoning underneath it.
This page is generated from the real, current implementation — every capability below is labeled Available (shipped and working today), Beta (real, still maturing), or Planned (roadmap, not yet built). Nothing here is aspirational marketing copy.
The core loop
Underneath every feature, MOS runs the same loop. It is deterministic code that calls an LLM for reasoning — never an LLM that is trusted to run the loop itself.
The model doing the reasoning is a swappable component. Operational state — goals, plans, actions, results, memory, permissions — lives in MOS, not in the model's context window, and survives switching providers or upgrading models entirely.
Architecture overview
The user talks to MOS — one conversational interface. Behind it, an Attention layer decides what deserves surfacing now versus what stays quiet, and routes work into goals, memory retrieval and internal agent roles.
Attention / routing layer
Vault, docs, operational records
Web / YouTube / X / Audio / Document Intelligence
Core capabilities
Goals
A goal is compiled from plain language into a structured metric, target and deadline, then broken into milestones and actions. Progress is a deterministic calculation from actual action state — never an LLM's guess at a percentage. Replanning is triggered by named, explicit conditions (repeated failures, a blocked dependency, an exhausted plan, an approaching deadline) — never on every step.
Autonomous execution — the Goal Runner
A bounded, resumable pass selects the next eligible action deterministically (dependency-aware, priority-ordered), respects your autonomy mode (advisory, assisted, autonomous) and spend policy, and either executes it through an internal agent role or hands it back to you. Execution is scheduler-driven — a stuck or crashed run is detected and surfaced, never blindly retried.
MOS Attention
A single, ranked view of what actually needs you: pending approvals, blocked work, an agent result material enough to matter, a goal drifting off track. Routine internal steps (a tool call, a search) never interrupt — they stay in the Activity log. Priority is deterministic: critical, needs a decision, an important result, informational, routine.
Persistent memory
Three layers, one retrieval surface: explicit facts you asked MOS to remember, documents in your Personal Vault, and MOS's own operational history (goals, actions, results). memory.search finds it; memory.ask answers a question grounded in what it finds, and says so plainly when nothing matches rather than guessing. See Memory & Vault below for the honest retrieval-strategy detail.
Intelligence sources
MOS separates two different things that are easy to conflate:
AI Engines: Anthropic · OpenAI · Gemini · xAI (Grok) · Mistral · Groq · DeepSeek · OpenRouter · Any OpenAI-compatible endpoint (incl. self-hosted Ollama)
| Provider | Capability | Status |
|---|---|---|
| Gemini | YouTube Intelligence — analyzes actual video/audio content, not just a transcript | Available |
| Tavily | Web Intelligence — deep web research & synthesis | Available |
| Brave Search | Web Intelligence — current web/news search with recency filtering | Available |
| Perplexity | Web Intelligence — cited, synthesized web answers | Available |
| xAI (Grok) | X Intelligence — live X/Twitter search via Grok's own X Search tool | Available |
| Groq | Audio Intelligence — Whisper transcription & translation, timestamped | Available |
| Mistral | Document Intelligence — real OCR, structured extraction, table-aware | Available |
BYOK & model independence
MOS does not depend on one proprietary AI vendor. Connect Anthropic, OpenAI, Gemini, xAI, Mistral, Groq, DeepSeek or OpenRouter with your own key — including OpenRouter via a real OAuth (PKCE) flow, no key-pasting required — or point a custom OpenAI-compatible endpoint at a self-hosted model (Ollama and similar work today). Every stored credential is encrypted at rest; provider usage belongs to you and is billed by that provider, not by Moseisley.
Moseisley is BYOK-only — it never pays LLM inference costs for you, on any plan. Every account connects its own OpenRouter key at signup; the MOS subscription covers the MOS platform, not bundled AI usage — see Pricing. Moseisley hosting its own local LLM/GPU compute for you is planned, not available yet.
Memory & the Personal Vault
The pipeline, precisely:
The original file is always kept — indexing failure never loses it. Retrieval is hybrid: PostgreSQL native lexical full-text search (indexed, not a LIKE scan) always runs, joined by semantic search — cosine similarity over embeddings from your own connected provider (BYOK) — whenever you've embedded content. Both are scoped to your tenant, and every result says plainly which one found it, never claiming semantic understanding it didn't use. There is no separate vector database: embeddings are ranked in plain PostgreSQL, so self-hosting never needs one either. Documents are deduplicated by checksum within your own account only. Cold-storage archival (e.g. Glacier) is a planned tier, not a live integration — nothing is silently moved anywhere colder today.
Agents & autonomy
You interact with MOS, not with a roster of agents you have to coordinate yourself. MOS delegates bounded, tracked jobs internally:
Internal roles available for delegation today:
- ·Strategist — Strategy and prioritization
- ·Challenger — Attempts to prove the current strategy wrong
- ·X-Ray — Analyzes historical and current reality for missed money/time
- ·Radar — External market intelligence, swept on a schedule
- ·Auditor — Verifies predictions against outcomes; keeps calibration honest
Every delegation is bounded (a fixed cap per turn), assignment is explicit (a crew role or the user — never an invented agent), and execution state is one of a small closed set: pending, in progress, blocked, completed, cancelled. A goal can be paused and resumed; a blocked action carries an explicit, honest reason; a stale or crashed run is detected and surfaced for review rather than silently retried. Internal delegation detail (which role ran what, in what order) is visible in Activity/X-Ray for anyone who wants it — MOS doesn't hide it, it just doesn't lead with it.
Safety & control
Give agents permission and budgets, not unrestricted authority. Concretely:
- ·Five kill switches — including one master Emergency Stop — checked at execution boundaries, not by asking the model politely.
- ·Tool Broker — the single, deterministic gate between an agent and any real integration; agents never hold raw credentials.
- ·Policy Engine — evaluates every tool call against your configured permissions before it runs.
- ·Spend policy & approvals — free-only, paid-allowed, or ask-before-spending; a real approval request blocks the action until you decide.
- ·Sandbox execution layer — used for untrusted code paths (e.g. the Dev Agent's own patches); denies execution by default rather than silently downgrading isolation.
- ·Tenant isolation — PostgreSQL Row-Level Security on tenant-owned data, enforced at the database layer, not only in application code.
- ·Encrypted credentials (AES-256-GCM) and an append-only audit Ledger for every meaningful state change.
Deliberately not detailed here: exact sandbox runtimes, infrastructure topology, or anything that would help someone probe for a weakness rather than understand the architecture.
Spending & economics
BYOK usage is never billed by Moseisley — your provider bills you directly for your own key's usage, and MOS's spend policy only controls when that spend is allowed to happen (never silently, never past your chosen limit). Hosted MOS is one flat monthly subscription for the MOS platform ($19/month after a 7-day free trial, still BYOK); the open-source code can also be self-hosted — see Pricing for current numbers.
Interfaces
Planned, not yet built: a dedicated phone app, glasses, wearables, and smart-home device adapters. Full-duplex realtime voice conversation is also planned — today is transcribe-then-send, not an always-listening line.
MCP & agent-to-agent
MOS includes a real MCP client: connect any Model Context Protocol server as a data/tool source, over the standard streamable-HTTP JSON-RPC transport, with the same policy and kill-switch enforcement as every other integration. What it returns is treated as untrusted content, like any external result.
Developer philosophy
- ·Open where possible — self-hostable, fair-code licensed
- ·Provider-independent — no proprietary model dependency
- ·Bring your own key — your usage, your cost, your data
- ·Standards-based — MCP over a bespoke protocol, where a standard exists
- ·Providers are replaceable — swap the model without losing state
- ·Persistent state outlives the model — goals, memory and permissions live in MOS
Available, Beta & Planned
Available
Beta
Planned
Get started
- 1Create an account. Hosted at moseisley.sh, or self-host from the GitHub repo.
- 2Connect an AI Engine. OpenRouter (OAuth or key), or any provider you already have a key for — including a local Ollama endpoint.
- 3Optionally connect Intelligence Sources. Tavily/Brave for web, Gemini for YouTube, Grok for X, Groq for audio, Mistral for documents — each is opt-in and BYOK.
- 4Create a Goal. Describe it in plain language. MOS compiles it into a metric, target and deadline.
- 5Let MOS plan and act. It generates milestones and actions, assigns what it can to the crew, and asks before anything that costs money.
- 6Add personal memory. Attach a document in chat and say "remember this" — MOS indexes it into your Personal Vault for later retrieval.
Prefer to self-host? The README has a five-minute Docker quickstart and an operations guide covering AI modes and deployment in detail.