DeepTutor: Lifelong Personalized Tutoring
Features · Get Started · Explore · CLI · Ecosystem · Community
🤝 We welcome any kinds of contributing! Vote on roadmap items or propose new ones at
Roadmap, and see our Contributing Guide for branching strategy, coding standards, and how to get started.
📦 Releases
[2026.9.27] v1.6.12 — Workspace KB moves, Kiwix archives, source figures, Task Board, German UI, and chat recovery.
[2026.9.24] v1.6.11 — French and Ukrainian interfaces, reading folders with in-chat selection actions, figures sent to vision models, paged Office previews, and self-syncing knowledge bases.
[2026.9.22] v1.6.10 — Native LightRAG role models, a published index that records and enforces what built it, PDF attachments that follow your parsing engine, visible truncation, and unfiltered provider choices.
[2026.9.21] v1.6.9 — Folder-based learning workspaces, daily practice, redesigned Settings, clearer streaming conversations, persistent usage accounting, and recoverable archives with explicit permanent deletion.
Past releases (more than 1 week ago)
[2026.9.14] v1.6.8 — A recycle bin for deleted chats, search across your whole conversation history, a tool that looks past your knowledge base, and a sweep of fixes for quiet failures.
[2026.9.11] v1.6.7 — A fix release: books that arrived as one empty chapter, quizzes that produced nothing, the model's scratchpad in the text, formulas printed raw, and cards you could not submit.
[2026.9.8] v1.6.6 — A fix release: answers that could not submit, a copy button that lied, connected knowledge bases for partners, Codex sign-in inside Docker, and a 100 KB lighter home route.
[2026.9.6] v1.6.5 — A content workspace you point at any folder, one
exectool for every language, Mastery Path modes that gate its tools, and Settings that grades readiness.
[2026.9.3] v1.6.4 — Faster isolated runtimes, controllable Book generation, source-complete Mastery paths and Chat hand-offs, durable Reading, unified activity UI, recoverable sessions, and explicit per-model API capabilities.
[2026.9.2] v1.6.3 — Breaking front/back-end refactor, strict canonical routes and recoverable streams, plus learner/guardian accounts, grounded Reading, WeKnora, broader parsing, Python 3.14, and DashScope media.
[2026.8.31] v1.6.2 — Immersive YouTube learning, a plugin-driven Visualize catalog, three new agent harnesses, safer reading citations, multi-format MinerU, live Partner channel status, and guided updates.
[2026.8.30] v1.6.1 — One vendor key linked to every service it serves, a task model for background work, Settings as a searchable navigator, a sidebar you arrange, and first-party LightRAG.
[2026.8.27] v1.6.0 — Faithful EPUB reading and annotations, Courses with Little Tutor and Ask Questions, bounded web-source sync, shared Books with private learning state, and Serply/native search.
[2026.8.25] v1.5.17 — Partners each member owns with private conversations and linkable chat accounts, GitHub repos as a knowledge source, Antigravity CLI, browser WeChat QR login, and
deeptutor doctor.
[2026.8.22] v1.5.16 — MarginNote 4 libraries you connect and its add-on fills, Book pages that turn again, and tool-call ids, embeddings and temperature limits that stop breaking behind a gateway.
[2026.8.20] v1.5.15 — PageIndex OSS you host yourself with reasoning retrieval, a question bank you can finally file into, third-party tool/capability plugins, and Apache Tika parsing.
[2026.8.19] v1.5.14 — Immersive Reading: a document open beside the thread, cited page by page; DeepTutor configures itself from chat; IMA libraries you browse and write to; a notebook console.
[2026.8.17] v1.5.13 — Books stream while they compile, track your progress, and export to Markdown; a cost estimate before you approve a spine; and home starter suggestions drawn from memory.
[2026.8.13] v1.5.12 — Web search rebuilt with six new providers (Doubao, Bocha, Zhipu, Firecrawl, Qianfan, Aliyun IQS), a LiteParse parsing engine, MCP servers that reconnect on credential change, and CodeBuddy + OrcaRouter.
[2026.8.10] v1.5.11 — Prose around a DSML tool call stops vanishing, a truncated reply continues instead of ending, live memory usage in Settings, and LightRAG indexing off the event loop.
[2026.8.7] v1.5.10 — Every account signs in to its own Codex, model output language becomes its own setting, empty tool calls are rejected instead of retried, and uploads stop blocking the loop.
[2026.8.4] v1.5.9 — Gemini Embedding 2 on its native endpoint, a per-model reasoning effort control, a Novita AI gateway, retrieval roles for queries, and Compose deployments that keep all of
data/.
[2026.8.2] v1.5.8 — Memory: a real heap ceiling for the dev server, source installs serve a production build, bounded LLM client and index caches, and a keep-alive fix for stray 500s.
[2026.7.31] v1.5.7 — A per-account MCP Services store, 101 CLI Apps the tutor can run, credentials moved out of the sandbox's reach, and a mobile layout.
[2026.7.29] v1.5.6 — Remote Codex sign-in completes behind an SSH tunnel, generated files get their own card in Activity, non-English languages stop collapsing to Chinese, and book creation no longer times out.
[2026.7.26] v1.5.5 — Sign in with your ChatGPT plan via OpenAI Codex OAuth, an Eden AI provider, knowledge bases that report what they hold, traceable
ragcitations, and GraphRAG indexing without a workaround.
[2026.7.24] v1.5.4 — Maintenance sweep: the post-answer "generating" stall is gone, IM partners render Markdown tables faithfully, LLM JSON parsing is sturdier, plus quiz, create-KB form, and Math Animator fixes.
[2026.7.24] v1.5.3 — Themeable code blocks, four more coding CLIs in My Agents (Gemini, Kimi, opencode, MiMo), an Atlas Cloud LLM provider, and a broad chat, memory, embedding, and parsing reliability sweep.
[2026.7.19] v1.5.2 — Configurable chat attachment limits, PageIndex retrieval that reasons across your documents via agentic tool calls, broader Anthropic/OpenAI model support, and steadier Book, Knowledge Base, and chat UI.
[2026.7.9] v1.5.1 — Remove a single failed document from a knowledge base — even one stuck in an error state — instead of deleting and rebuilding the whole base.
[2026.7.4] v1.5.0 — LlamaIndex ingestion now honors your Document Parsing engine with multimodal image extraction, Partner & Soul ids stay URL-safe for non-Latin names, and optional RAG extras install cleanly on Python 3.14+.
[2026.6.30] v1.4.15 — A native Mattermost channel for Partners, plus fixes so Guided Learning multiple-choice questions grade correctly and a configured zero chunk overlap is honored.
[2026.6.29] v1.4.14 — Click an assigned partner to chat in one step, Deep Research flags partial reports, LightRAG indexes without MinerU, FAISS handles non-ASCII paths, and PocketBase sessions are isolated per user.
[2026.6.27] v1.4.13 — Partners support non-Latin names and become assignable to users, logos render after login (#599), tiny knowledge bases retrieve reliably, and containers start cleanly under rootless Podman.
[2026.6.24] v1.4.12 — A new LightRAG Server retrieval engine, a lightweight PyMuPDF4LLM parsing engine, and a FAISS vector backend that makes large knowledge-base retrieval dramatically faster.
[2026.6.23] v1.4.11 — Native tool calling on every cloud OpenAI-compatible provider, a redesigned admin Users page, LaTeX in quiz options, an honest session-loading spinner, and configurable container host binding.
[2026.6.21] v1.4.10 — A self-service Profile page with avatars, a rootless-ready container guide with a single-port request-time proxy, and deny-by-default MCP tools for non-admin users.
[2026.6.19] v1.4.9 — Settings polish: Search shows only the fields your provider needs, connection profiles can be renamed and auto-named by provider, and graded Mastery Path questions flow into your Question Bank.
[2026.6.18] v1.4.8 — Connect your own Partners under My Agents and consult them live in chat — answering through their own persona, library and skills — each with its own private memory.
[2026.6.18] v1.4.7 — Connect your local Claude Code / Codex and consult it live mid-turn, My Agents graduates to a top-level
/agents, and Partner conversations gain branch / resume / delete with a replayable trace.
[2026.6.17] v1.4.6 — Four-surface consolidation: a Space learning dashboard with importable My Agents and top-level Memory, a Knowledge Center with GraphRAG / PageIndex / LightRAG / linked-KB / Obsidian, opened-up Settings, and per-model capability gating.
[2026.6.14] v1.4.5 — Guided Learning rebuilt on the chat agent loop with a hard per-type mastery gate and a
/learningdashboard, a new loop-plugin framework, plus Markdown export / save-to-notebook for Partner conversations.
[2026.6.13] v1.4.4 — Install community skills from ClawHub with
deeptutor skill installbehind a security gate, plus real in-browser DOCX/XLSX previews for knowledge-base files.
[2026.6.12] v1.4.3 — TutorBot becomes Partners on a production-grade IM pipeline (15 channels, live streaming), Chat moves to a single agent loop, real per-user isolation, and a rebuilt Visualize.
[2026.5.28] v1.4.2 — Stability + polish: Gemini 2.5+ unblocked across Visualize and Chat, auth-routing fix (#485), smooth-streaming chat UX, a Recents sidebar, and Lemonade local-provider support.
[2026.5.27] v1.4.1 — Security + stability: TutorBot tool sandbox locked down, per-user resource isolation, multimodal image fallback, an HTTP/SSE API for TutorBots, and a v1.4.0 chat regression fix.
[2026.5.22] v1.4.0 — GA cut of v1.4: Auto Mode, three-layer Memory, agentic Deep Research / Solve / Question, LlamaIndex RAG refactor, Visualize/Animator merge, and restart-safe turn runtime.
[2026.5.21] v1.4.0-beta — Three-layer Memory workbench (L1/L2/L3), every chat capability rebuilt on a single agentic engine, LlamaIndex-only RAG, and a unified Settings + Capabilities surface.
[2026.5.10] v1.3.10 — Remote Docker CORS recovery,
DISABLE_SSL_VERIFYacross SDK providers, safer code-block citations, and optional Matrix E2EE add-on.
[2026.5.9] v1.3.9 — TutorBot Zulip and NVIDIA NIM support, safer thinking-model routing,
deeptutor start, sidebar tooltips, and session-store parity.
[2026.5.8] v1.3.8 — Optional multi-user deployments with isolated user workspaces, admin grants, auth routes, and scoped runtime access.
[2026.5.4] v1.3.7 — Thinking-model/provider fixes, visible Knowledge index history, and safer Co-Writer clear/template editing.
[2026.5.3] v1.3.6 — Catalog-based model selection for chat and TutorBot, safer RAG re-indexing, OpenAI Responses token-limit fixes, and Skills editor validation.
[2026.5.2] v1.3.5 — Smoother local launch settings, safer RAG queries, cleaner local embedding auth, and Settings dark-mode polish.
[2026.5.1] v1.3.4 — Book page chat persistence and rebuild flows, chat-to-book references, stronger language/reasoning handling, RAG document extraction hardening.
[2026.4.30] v1.3.3 — NVIDIA NIM + Gemini embedding support, unified Space context for chat history/skills/memory, session snapshots, RAG re-index resilience.
[2026.4.29] v1.3.2 — Transparent embedding endpoint URLs, RAG re-index resilience for invalid persisted vectors, memory cleanup for thinking-model output, Deep Solve runtime fix.
[2026.4.28] v1.3.1 — Stability: safer RAG routing & embedding validation, Docker persistence, IME-safe input, Windows/GBK robustness.
[2026.4.27] v1.3.0 — Versioned KB indexes with re-index workflow, rebuilt Knowledge workspace, embedding auto-discovery with new adapters, Space hub.
[2026.4.25] v1.2.5 — Persistent chat attachments with file-preview drawer, attachment-aware capability pipelines, TutorBot Markdown export.
[2026.4.25] v1.2.4 — Text/code/SVG attachments, one-command Setup Tour, Markdown chat export, compact KB management UI.
[2026.4.24] v1.2.3 — Document attachments (PDF/DOCX/XLSX/PPTX), reasoning thinking-block display, Soul template editor, Co-Writer save-to-notebook.
[2026.4.22] v1.2.2 — User-authored Skills system, chat input performance overhaul, TutorBot auto-start, Book Library UI, visualization fullscreen.
[2026.4.21] v1.2.1 — Per-stage token limits, Regenerate response across all entry points, RAG & Gemma compatibility fixes.
[2026.4.20] v1.2.0 — Book Engine "living book" compiler, multi-document Co-Writer, interactive HTML visualizations, Question Bank @-mention.
[2026.4.18] v1.1.2 — Schema-driven Channels tab, RAG single-pipeline consolidation, externalized chat prompts.
[2026.4.17] v1.1.1 — Universal "Answer now", Co-Writer scroll sync, unified settings panel, streaming Stop button.
[2026.4.15] v1.1.0 — LaTeX block math overhaul, LLM diagnostic probe, Docker + local LLM guidance.
[2026.4.14] v1.1.0-beta — Bookmarkable sessions, Snow theme, WebSocket heartbeat & auto-reconnect, embedding registry overhaul.
[2026.4.13] v1.0.3 — Question Notebook with bookmarks & categories, Mermaid in Visualize, embedding mismatch detection, Qwen/vLLM compatibility, LM Studio & llama.cpp support, and Glass theme.
[2026.4.11] v1.0.2 — Search consolidation with SearXNG fallback, provider switch fix, and frontend resource leak fixes.
[2026.4.10] v1.0.1 — Visualize capability (Chart.js/SVG), quiz duplicate prevention, and o4-mini model support.
[2026.4.10] v1.0.0-beta.4 — Embedding progress tracking with rate-limit retry, cross-platform dependency fixes, and MIME validation fix.
[2026.4.8] v1.0.0-beta.3 — Native OpenAI/Anthropic SDK (drop litellm), Windows Math Animator support, robust JSON parsing, and full Chinese i18n.
[2026.4.7] v1.0.0-beta.2 — Hot settings reload, MinerU nested output, WebSocket fix, and Python 3.11+ minimum.
[2026.4.4] v1.0.0-beta.1 — Agent-native architecture rewrite (~200k lines): Tools + Capabilities plugin model, CLI & SDK, TutorBot, Co-Writer, Guided Learning, and persistent memory.
[2026.1.23] v0.6.0 — Session persistence, incremental document upload, flexible RAG pipeline import, and full Chinese localization.
[2026.1.18] v0.5.2 — Docling support for RAG-Anything, logging system optimization, and bug fixes.
[2026.1.15] v0.5.0 — Unified service configuration, RAG pipeline selection per knowledge base, question generation overhaul, and sidebar customization.
[2026.1.9] v0.4.0 — Multi-provider LLM & embedding support, new home page, RAG module decoupling, and environment variable refactor.
[2026.1.5] v0.3.0 — Unified PromptManager architecture, GitHub Actions CI/CD, and pre-built Docker images on GHCR.
[2026.1.2] v0.2.0 — Docker deployment, Next.js 16 & React 19 upgrade, WebSocket security hardening, and critical vulnerability fixes.
✨ v1.6.12 is live.
pip install -U deeptutorpicks up the latest stable release.
📰 News
- 2026-09-20 🎉 40k stars in 9 months! We'll keep expanding DeepTutor's learning ecosystem.
- 2026-05-22 🌐 Official docs site live at deeptutor.info — guides, references, and capability tours in one place.
- 2026-04-19 🎉 20k stars in 111 days! Thank you for the support toward truly personalized, intelligent tutoring.
- 2026-04-10 📄 Our paper is live on arXiv — read the preprint for the design and ideas behind DeepTutor.
- 2026-02-06 🚀 10k stars in just 39 days! A huge thank you to our incredible community.
- 2026-01-01 🎊 Happy New Year! Join our Discord, WeChat, or Discussions — let's shape DeepTutor together.
- 2025-12-29 🎓 DeepTutor is officially released!
✨ Key Features
DeepTutor is an agent-native learning workspace that connects tutoring, problem solving, quiz generation, research, visualization, and mastery practice in one extensible system.
- One runtime for every mode — Chat, Ask Questions, Quiz, Research, Visualize, Solve, Course Study, Mastery Path, Immersive Reading, and Immersive Watching share one capability runtime and session context while keeping purpose-built loops and pipelines.
- Task Board — Track study tasks in To do, In progress, and Done, with notes, drag-and-drop or keyboard-accessible move buttons, and an archive you can restore from. Cards stay in the current workspace and follow the existing appearance and language settings; no model configuration is required.
- Connected learning context — Knowledge bases, books, Co-Writer drafts, notebooks, question banks, personas, and Memory can be reused across the workflows that support them, subject to account grants and learning policies.
- Immersive video learning — paste a YouTube link for privacy-enhanced native playback, synchronized captions, timestamp-grounded tutoring, and resumable progress; administrators can switch playback to a self-hosted Invidious instance without rebuilding materials.
- Subagents and Partners — from Chat, consult a live agent harness (Claude Code, Codex, Grok CLI, Antigravity, Kimi, opencode, MiMo, Hermes, OpenClaw, or DeepSeek) or a Partner, import past conversations, and run persistent IM companions on the same brain.
- Multi-engine knowledge — versioned RAG libraries across LlamaIndex, PageIndex, GraphRAG, LightRAG, a remote LightRAG Server, a self-hosted WeKnora knowledge base, a Tencent IMA or MarginNote 4 library, a connected Kiwix ZIM archive, or a linked Obsidian vault, with pluggable document parsing. See native LightRAG role models for independent extraction, query and vision settings, default-only creation and confirmed rebuilds.
- Extensible tools and skills — built-in tools, MCP servers, CLI apps, image / video / voice generation models, and installable community skills from EduHub.
- Inspectable memory — L1 traces, L2 surface summaries, and L3 synthesis make personalization visible and editable; the Memory Graph links L2 facts to L1 evidence and L3 synthesis to contributing surfaces.
🚀 Get Started
DeepTutor ships four installation paths. They all share one runtime-home layout: private settings live in data/user/settings/ under the directory you launch from (or under DEEPTUTOR_HOME / deeptutor start --home if you set one explicitly). For the full app, the recommended flow is pick a runtime-home directory → install → deeptutor init → deeptutor start.
Content Workspace
The Content Workspace is separate from DeepTutor's private runtime home. It
is the folder agents may read, with generated files under
outputs/<capability>/<session>/<turn>/. Custom workspaces isolate conversations,
learning materials, progress, and caches in a private .deeptutor/data/ tree
that file tools cannot browse. Settings, credentials, and Memory stay shared
at the account level.
Without configuration, the content workspace is
<runtime-home>/data/user/workspace. Local PyPI, CLI, and source installs can
choose folders in Settings → Workspaces; set the default folder with:
deeptutor workspace show
deeptutor workspace set /absolute/path/to/my-folder
deeptutor workspace resetCapabilities inspect their selected workspace through the built-in workspace
tools. The model only receives relative paths such as outputs/...; when it
uses workspace_present, the UI renders an authenticated, openable snapshot.
The same exact relative path also works in a normal Markdown link or image.
Changing the source file later does not change an already presented snapshot.
Learning Space manages the resource library. In Settings → Workspaces, assign Skills, MCP services and knowledge bases to each workspace, or retain its existing access rules. Existing workspaces keep their current access until a selection is saved. Assignments reference the original resources without copying credentials or knowledge indexes; workspace-specific skills can override shared versions. See workspace resource assignments.
Execution is read-only outside outputs/. Copying a generated file elsewhere
in the content workspace requires an explicit Allow once confirmation for
that exact source and destination. A system sandbox or the Docker runner
enforces this boundary when available; local restricted-subprocess fallback is
shown as best effort in Workspace settings.
Option 1 — Install From PyPI · full local Web app + CLI, no clone required
Full local Web app + CLI, no clone required. Needs Python 3.11–3.14 and a Node.js 20+ runtime on PATH (the packaged Next.js standalone server is spawned by deeptutor start).
mkdir -p my-deeptutor && cd my-deeptutor
pip install -U deeptutor
deeptutor init # prompts for ports + LLM provider + optional embedding/search
deeptutor start # starts backend + frontend; keep the terminal opendeeptutor init prompts for backend port (default 8001), frontend port (default 3782), LLM provider / base URL / API key / model, an optional embedding provider for Knowledge Base / RAG, and an optional search provider for Web Search.
After deeptutor start, open the frontend URL printed in the terminal — by default http://127.0.0.1:3782. Press Ctrl+C in that terminal to stop both backend and frontend. Skipping deeptutor init is fine for a quick trial; the app boots with default ports and empty model settings, configure them later in Settings → Providers and Language models.
Browser microphone transcription: OpenAI-compatible STT adapters forward browser audio to the provider without local conversion. The native DashScope and Volcengine STT adapters convert browser WebM/Opus to 16 kHz WAV and require an ffmpeg executable on DeepTutor's PATH. A canonical 16 kHz mono PCM WAV bypasses this conversion. For Windows PyPI installations using either native adapter, install FFmpeg, add its bin directory to the service's PATH, then restart DeepTutor. Conversion failures appear below the chat input.
Option 2 — Install From Source · develop against a checkout
For development against a checkout. Use Python 3.11–3.14 and Node.js 22 LTS to match CI and Docker.
git clone https://github.com/HKUDS/DeepTutor.git
cd DeepTutor
# Create a venv (macOS/Linux). Windows PowerShell:
# py -3.11 -m venv .venv ; .\.venv\Scripts\Activate.ps1
python3 -m venv .venv && source .venv/bin/activate
python -m pip install --upgrade pip
# Install backend + frontend deps
python -m pip install -e .
( cd web && npm ci --legacy-peer-deps )
deeptutor init
deeptutor start --devdeeptutor start builds the local web/ frontend for production once and reuses it; --dev runs Next.js with HMR. Config layout, ports, and Ctrl+C match Option 1.
Conda environment (instead of venv)
conda create -n deeptutor python=3.11
conda activate deeptutor
python -m pip install --upgrade pipOptional install extras — RAG engines / dev / partners / matrix / math-animator
pip install -e ".[rag-lightrag]" # Built-in LightRAG engine (exact supported SDK)
pip install -e ".[graphrag]" # Microsoft GraphRAG engine (Python 3.11–3.13)
pip install -e ".[dev]" # tests/lint tools
pip install -e ".[partners]" # Partner IM channel SDKs
pip install -e ".[video-learning]" # compatibility extra; captions ship in the full/CLI installs
pip install -e ".[matrix]" # Matrix channel without E2EE/libolm
pip install -e ".[matrix-e2e]" # Matrix E2EE; requires libolm
pip install -e ".[math-animator]" # Manim addon; requires LaTeX/ffmpeg/system libsFrontend dependency tweaks & dev-server troubleshooting
Changing frontend dependencies: run npm install --legacy-peer-deps to refresh web/package-lock.json, then commit both web/package.json and web/package-lock.json.
Stuck dev server: if deeptutor start --dev reports an existing frontend that isn't responding, stop the PID it prints. If no Next.js process is actually running, the lock files are stale — remove them and retry:
rm -f web/.next/dev/lock web/.next/lock
deeptutor start --devOption 3 — Docker · one self-contained container
One container for the full Web app. Images on GitHub Container Registry:
ghcr.io/hkuds/deeptutor:latest— latest stable releaseghcr.io/hkuds/deeptutor:<version>— exact release without the leadingv(for example:1.6.3); pre-releases receive only their version tag
See CONTAINERIZATION.md for podman/rootless/read-only-rootfs deployments and the full per-installation guide.
docker run --rm --name deeptutor \
-p 127.0.0.1:3782:3782 \
-v deeptutor-data:/app/data \
ghcr.io/hkuds/deeptutor:latestTo choose a host content folder at container startup, mount it at the stable container path and lock DeepTutor to that path:
mkdir -p "$PWD/deeptutor-workspace/outputs"
docker run --rm --name deeptutor \
-p 127.0.0.1:3782:3782 \
-v deeptutor-data:/app/data \
-v "$PWD/deeptutor-workspace:/workspace" \
-e DEEPTUTOR_WORKSPACE_ROOT=/workspace \
-e DEEPTUTOR_WORKSPACE_ALLOWED_ROOTS=/workspace \
ghcr.io/hkuds/deeptutor:latestFor Compose, set DEEPTUTOR_WORKSPACE_HOST=/absolute/host/folder before
running python scripts/docker_compose.py up -d. When omitted it uses
./data/user/workspace. Docker paths are selected at startup and therefore
appear locked in the Web settings page.
Only
3782needs to be published. The browser talks exclusively to the frontend origin; the Next.js middleware (web/proxy.ts) forwards/api/*and/ws/*to the FastAPI backend inside the container. Publishing8001(-p 127.0.0.1:8001:8001) is optional — handy only for hitting the API directly with curl or scripts.
Open http://127.0.0.1:3782. The container creates /app/data/user/settings/*.json on first boot; configure model providers from the Web Settings page. Config, API keys, logs, the default Content Workspace, memory, and knowledge bases persist in the deeptutor-data volume. A separately mounted Content Workspace persists at its host path instead. Optional extras belong on the deployment, not in a shell: set DEEPTUTOR_EXTRAS (and DEEPTUTOR_APT_PACKAGES for system libraries) and every container started from it re-applies them, where a docker exec … pip install would be lost at the next compose down.
- Different host ports: change the left side of each
-p host:containermapping (e.g.-p 127.0.0.1:8088:3782). If you change container-side ports in/app/data/user/settings/system.json, restart and update the right side of each mapping to match. - Detached: add
-d, thendocker logs -f deeptutorto follow,docker stop deeptutorto stop,docker rm deeptutorbefore reusing the name. Thedeeptutor-datavolume keeps private runtime data and the default Content Workspace across restarts; a separately mounted Content Workspace persists at its host path.
Remote Docker / reverse proxy: the browser only talks to the frontend
origin (:3782); the in-container Next.js middleware forwards /api/* and
/ws/* to the backend server-side. For the common single-container case you
don't configure an API base at all — just point your reverse proxy / TLS
terminator at :3782. You only need an API base for a split deployment
(backend in a separate container/host): set next_public_api_base in
data/user/settings/system.json to the in-network address the frontend server
uses to reach the backend (it's read server-side, never sent to the browser).
{
"next_public_api_base": "http://backend:8001"
}next_public_api_base_external (and its alias public_api_base) are accepted as
lower-precedence fallbacks. CORS uses frontend origins, not API URLs. With
auth disabled, DeepTutor permits normal HTTP/HTTPS browser origins by default.
With auth enabled, add exact frontend origins:
{
"cors_origins": ["https://deeptutor.example.com"]
}Connecting to Ollama / LM Studio / llama.cpp / vLLM / Lemonade on the host
Inside Docker, localhost is the container itself, not your host machine. To reach a model service running on the host, use the host gateway (recommended):
docker run --rm --name deeptutor \
-p 127.0.0.1:3782:3782 -p 127.0.0.1:8001:8001 \
--add-host=host.docker.internal:host-gateway \
-v deeptutor-data:/app/data \
ghcr.io/hkuds/deeptutor:latestThen in Settings → Providers, point the provider Base URL at host.docker.internal:
- Ollama LLM:
http://host.docker.internal:11434/v1 - Ollama embedding:
http://host.docker.internal:11434/api/embed - LM Studio:
http://host.docker.internal:1234/v1 - llama.cpp:
http://host.docker.internal:8080/v1 - Lemonade:
http://host.docker.internal:13305/api/v1
Docker Desktop (macOS/Windows) usually resolves host.docker.internal without --add-host. On Linux, the flag is the portable way to create that hostname on modern Docker Engine.
Linux alternative — host networking: add --network=host and drop the -p flags. The container shares the host network directly, so open http://127.0.0.1:3782 (or the frontend_port in system.json), and host services can be reached with normal localhost URLs like http://127.0.0.1:11434/v1. Note that host networking exposes container ports directly on the host and may conflict with existing services — to keep them on loopback, set BACKEND_HOST=127.0.0.1 and FRONTEND_HOST=127.0.0.1 (see CONTAINERIZATION.md).
Option 4 — CLI Only · no Web UI, from a source checkout
When you don't need the Web UI. The CLI-only package is installed from a source checkout, not from PyPI.
git clone https://github.com/HKUDS/DeepTutor.git
cd DeepTutor
# Create a venv (macOS/Linux). Windows PowerShell:
# py -3.11 -m venv .venv-cli ; .\.venv-cli\Scripts\Activate.ps1
python3 -m venv .venv-cli && source .venv-cli/bin/activate
python -m pip install --upgrade pip
python -m pip install -e ./packaging/deeptutor-cli
deeptutor init --cli
deeptutor chatdeeptutor init --cli shares the same data/user/settings/ layout as the full app but skips the backend/frontend port prompts. It still offers the Embedding and Search selectors (choose Skip when you do not need them), writes the key runtime files (system.json, auth.json, integrations.json, interface.json, model_catalog.json, main.yaml, agents.yaml), and prompts for the active LLM provider and model.
Common commands
deeptutor chat # interactive REPL
deeptutor chat --capability deep_solve --tool rag --kb my-kb
deeptutor run chat "Explain Fourier transform"
deeptutor run deep_solve "Solve x^2 = 4" --tool rag --kb my-kb
deeptutor kb create my-kb --doc textbook.pdf
deeptutor memory show
deeptutor config showThe local deeptutor-cli install ships no Web assets or server dependencies. Keep the source checkout around — the editable install points to it. To add the Web app later, install the PyPI package (Option 1) and run deeptutor init + deeptutor start from the same workspace.
Code Execution Sandbox (office skills) · running model-generated code for docx / pdf / pptx / xlsx
The built-in office skills — docx / pdf / pptx / xlsx — work by having the
model write a short Python script (python-docx, reportlab, openpyxl, …),
run it through the single exec tool, and present the saved workspace file.
Those tools mount whenever a sandbox backend is active. DeepTutor selects the
strongest configured backend in this order:
- Runner sidecar:
DEEPTUTOR_SANDBOX_RUNNER_URLroutes execution to the hardened, least-privileged service fromDockerfile.runner. - Linux bubblewrap: when available,
bwrapisolates the process and files. - Restricted subprocess fallback: local and single-container installs use this only when allowed; under Docker the container remains another boundary.
The sandbox_allow_subprocess setting in data/user/settings/system.json
(default true) controls only the last fallback. Set it to false (or export
DEEPTUTOR_SANDBOX_ALLOW_SUBPROCESS=0) to refuse subprocess execution when no
runner or bwrap backend is available; it does not disable those stronger
backends.
Configuration reference — config files under data/user/settings/ (JSON/YAML)
Everything under data/user/settings/ is plain JSON/YAML. The Settings page is the recommended editor; workspace registrations live separately in data/user/.runtime/workspaces.sqlite3.
| File | Purpose |
|---|---|
model_catalog.json |
Provider connections plus LLM, task, embedding, search, TTS, STT, image, and video profiles, credentials, and active selections |
system.json |
Backend/frontend ports, public API base, CORS, SSL verification, attachment directory and upload/extraction limits |
auth.json |
Optional auth toggle, username, password hash, token/cookie settings |
integrations.json |
Optional PocketBase and sidecar integration settings |
interface.json |
UI and model output language / theme / sidebar preferences |
document_parsing.json |
Parsing engine selection, remote endpoints, and engine-specific options |
video_learning.json |
Default YouTube/Invidious playback provider, Invidious origins, and optional transcript adapter |
main.yaml |
Runtime behavior defaults and path injection |
agents.yaml |
Capability/tool temperature and token settings |
Web Search references are filtered by default: only public http/https URLs
without embedded credentials or unusual ports are surfaced. Deployments can add
an education-focused domain policy in data/user/settings/system.json:
{
"web_search_source_filtering": {
"enabled": true,
"blocked_domains": ["spam.example"],
"trusted_domains": ["edu.cn", "arxiv.org"]
}
}When trusted_domains is non-empty, references are limited to those domains
and their subdomains; blocked_domains always takes precedence.
Project-root .env is not read as an application config file. For a minimal model setup, save a Base URL and API key in Settings → Providers, then add and select an LLM in Language models. Add an embedding profile only if you plan to use Knowledge Base / RAG features.
LLM and task-model profiles expose an API format setting when their provider
supports a choice. Keep Auto for normal routing and fallback, or choose
OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages; forced
Responses remains fail-closed. The persisted field is api_format (auto,
openai_chat, openai_responses, or anthropic); wire_api is derived
compatibility state. Per-model Auto / Supported / Not supported overrides
cover tool calling, image input, JSON output, and reasoning controls.
Uninstall and cleanup
DeepTutor separates its installed code, private runtime home, and optional
Content Workspace. By default, the runtime home is the directory where you run
deeptutor init / deeptutor start; --home PATH or DEEPTUTOR_HOME
overrides it. Private application state is the data directory inside that
home, so the startup banner line beginning with Workspace: identifies that
runtime location. If Settings → Workspace points to another folder, back up
or remove that content folder separately; it is intentionally not erased by
uninstalling DeepTutor.
-
Stop the app. Press
Ctrl+Cin the terminal runningdeeptutor start, or rundeeptutor stop [--home PATH]for a launcher started with--detach; stop any running Partner and detached Docker containers before deleting data. -
Remove runtime data only if you also want to erase all local state. This includes settings and API keys, chat history, sessions, Memory, Notebooks, Books, Reading state, Skills, Partners state, logs, Knowledge Bases, parse caches, generated artifacts, and the packaged frontend runtime cache.
First copy the exact
Workspace:path from the startup banner and verify that itsdatachild is the intended DeepTutor data directory. Back it up if anything may be needed later, then move that exact directory to your operating system's Trash/Recycle Bin. Do not run a recursive deletion command against a relative path or an unresolved environment variable. -
Remove the installed package. Use the command that matches the distribution:
python -m pip uninstall deeptutor python -m pip uninstall deeptutor-cli
If the virtual environment was created only for DeepTutor, remove it through your environment manager. For a source install, deactivate the environment, leave the source directory, and run
git status --shortinside that exact checkout. Only move the checkout to Trash/Recycle Bin after confirming it contains no unrelated or uncommitted work. -
For the Docker path, inspect the exact container and named volume before removing them. Volume removal permanently erases the Docker-managed data:
docker ps -a --filter name=^/deeptutor$ docker volume inspect deeptutor-data docker rm -f deeptutor docker volume rm deeptutor-data
📖 Explore DeepTutor
Start with the main surfaces you will use day to day: Chat, Partners, My Agents, Co-Writer, Book, Knowledge Center, Learning Space, Memory, and Settings. The tour then covers Multi-User deployments for shared, isolated workspaces.
If an answer loses an earlier constraint, cites weak evidence, or disagrees with selected material, collect the diagnostics in REASONING_SAFETY_CHECKLIST.md before opening an issue.
Screenshot status: The overview is current for v1.6.5. The surface screenshots below remain v1.4.6 references while a versioned refresh is in progress. Use them to understand workflows, not as exact navigation.
💬 Chat — The Agent Loop You Actually Use
Chat is the default capability and where most work begins. A single thread can talk normally, call tools, ground itself in selected knowledge bases, read attachments, generate images, consult subagents, write notebook records, and continue with the same context across turns.
The loop is deliberately simple: the model thinks in rounds, calls tools when useful, observes the results, and finishes with a tool-free message. ask_user is special — instead of guessing, the agent can pause the turn, ask a structured clarifying question, and resume once you answer.
User-toggleable tools are brainstorm, web_search, paper_search, zotero_search, reason, and geogebra_analysis — plus imagegen and videogen once you configure the matching generation model. Contextual tools such as rag, kb_files, knowledge_frontier, read_source, read_memory, write_memory, read_skill, load_tools, exec, web_fetch, ask_user, list_notebook, write_note, question_bank, github, consult_subagent, workspace_list, workspace_read, workspace_search, workspace_present, and workspace_export mount automatically when the turn has the right context.
Context comes in two kinds: sticky session context (capability, workspace or course, tools, knowledge bases, persona, model, and Reading / Mastery state) persists across turns; one-time references (files, chat history, books, reading sections, notebooks, question bank, imported agents) come from the + menu for a single turn. The voice button only transcribes the current message.
Home keeps Chat, Ask Questions, Quiz, and Visualize one click away; Research for cited reports, Solve for worked reasoning, and Immersive Watching sit under More Capabilities. Personalized Learning groups Book, Mastery Path, Immersive Reading, Watching, and Practice; Reading adds verified citations, saved notes, source-grounded read-aloud / study guidance / vocabulary / quiz / translation actions, and notebook capture, while Course Study keeps its course-bound context.
🤝 Partner — Persistent Companions on the Same Brain
Partners are persistent companions with their own soul, model policy, library, memory, and channels. They are not a separate bot engine: every inbound web or IM message becomes a normal ChatOrchestrator turn inside a partner-scoped workspace. A partner is "a chat that has a personality and a phone number."
Each partner has a SOUL.md, model selection, channels, tool policy, and assigned library. Its library either copies knowledge bases, skills, and notebooks into data/partners/<id>/workspace/, or stays linked to an existing workspace's files and resources; soul, conversations, and Partner memory remain separate. Authenticated non-admin users keep private partner sessions and relationship memory while the partner reads their personal memory read-only; admin, group, and unbound traffic use the shared partner scope.
The channel layer is schema-driven and can connect to IM platforms such as Feishu, Telegram, Slack, Discord, DingTalk, QQ/NapCat, WeCom, WhatsApp, Zulip, Mattermost, Matrix, Mochat, and Microsoft Teams depending on installed extras and configured credentials. A partner can also be connected as a subagent and consulted from a normal chat turn — see My Agents below.
For faster setup, the Partner channel page can create a Feishu/Lark app or WeCom AI bot, or sign a personal WeChat account in, from a QR scan drawn in the browser rather than the server log. Feishu/Lark detects the account domain and saves the scanning user as the initial allowed sender. WeCom keeps an existing allowlist and otherwise defaults to all users who can reach the bot, with a visible open-access warning; the manual channel forms remain available if a provider's scan protocol changes.
🧑🚀 My Agents — Consult & Import Other Agents
My Agents turns other agents into context for DeepTutor, and does two distinct things. Connect a live agent — Claude Code, Codex, Grok CLI, Antigravity, Kimi, opencode, MiMo Code, Hermes Agent, OpenClaw, or DeepSeek Harness on your machine, a remote Hermes gateway, or one of your Partners — and consult it from inside a chat turn: DeepTutor actually runs the other agent and streams its work into the Activity panel via the consult_subagent tool. Select it and its round limit with the Agent chip, or filter the same connected-agent list with @; the choice stays attached to the session.
Connect Grok CLI. Install xAI's Grok CLI on the machine running the DeepTutor
backend, run grok login there, and verify that grok --help lists
--output-format streaming-json. Then open My Agents → Connect, select
Grok CLI, and choose a working directory. Detection checks the executable's
protocol support; it does not verify login or model access. The connector has
been exercised with Grok CLI 1.0.3; unrelated third-party commands also named
grok are not supported.
In Settings → Partners & Agents → Grok CLI, leave model and reasoning effort
empty to use the CLI defaults, or enter values supported by your account's
grok models output. System instructions are passed through --rules. The
default permission mode is dontAsk: Grok uses its existing rules and built-in
read-only handling, and denies operations that need approval. This is a CLI
permission policy, not a filesystem sandbox. Broader modes can be selected
explicitly in settings; advanced CLI flags remain available through
backends.grok.extra_args in the subagent settings API.
Grok uses its own authentication and session storage; DeepTutor does not copy its credentials. Follow-up consults resume the connection's session in the same working directory. Text and tool activity stream live; private thought payloads are omitted. Cross-session Grok memory is disabled by default. This connector supports text questions and CLI tools, not image forwarding or importing past Grok conversations. In Docker, install and authe
(README truncated)








