Sign inSign up

jncchds/abook

By jncchds

•Updated 4 days ago

Image
0

4.8K

jncchds/abook repository overview

⁠ABook — Agentic AI Book Writing

License: MIT Docker Hub .NET

A self-hosted web application that uses AI agents to collaboratively write books. Seven specialized agents — Story Bible, Characters, Plot Threads, Chapter Outlines, Writer, Checker, and Editor — work together under your direction, streaming their progress in real time and pausing to ask clarifying questions when needed.


⁠Features

⁠Writing pipeline
  • Seven specialized agents across two phases: a 4-phase Planner (Story Bible → Characters → Plot Threads → Chapter Outlines) and a per-chapter writing pipeline (Pre-write Check → Write → Check → Edit)
  • Checker → mechanical patch apply — the Checker flags continuity, grammar, repetition, and style issues as structured JSON patches; the Editor applies them mechanically (no LLM call) using indexed text matching with whitespace normalization and position hints
  • Full synopsis spine — every prior chapter's title and outline is injected into Writer and Editor messages so agents stay aware of the whole narrative and avoid recycling beats or re-introducing established characters
  • RAG context retrieval — Writer runs 3 targeted queries (characters, locations, plot threads) and Editor runs 4 (same + repeated-phrase detection) against pgvector embeddings
⁠Planning & guidance
  • Guided planning Q&A — before planning, the planner posts its clarifying questions to the book's Clarifications page and waits; answer each one there (or skip it), and submitting the last answer completes the round and planning carries on. Answers are kept as clarifications and added to the premise in every prompt, so nothing is lost if planning is interrupted, and the round isn't repeated once it's done. Edit, add, archive or restore them on the same page, or reopen the round to be asked again
  • Human-assisted generation — pauses after each planning phase, and after each chapter's mechanical fixes so you can edit the book and steer the creative rewrite; pending questions are restored after page refresh; supports Ctrl+Enter to submit
  • Flexible workflow controls — Plan Only, Write Book, Continue, Continue Planning, and individual per-chapter agent buttons; Stop cancels any running agent cleanly
  • Refresh-safe progress — reloading the page (or losing the connection) mid-run keeps the current step, the latest progress message and everything streamed so far; new text continues where it left off
  • Interrupted runs keep their output — if a Characters, Plot Threads, or Chapter Outlines run times out, drops, or is stopped part-way, everything that finished streaming is saved instead of discarded; run the phase again to pick up where it left off
  • Re-runs build on what exists — regenerating a planning phase sends the current characters, plot threads, or outlines back to the model to refine and extend, rather than starting from a blank page; archived items stay out of it and are never overwritten
⁠Content management
  • Story Bible, Characters & Plot Threads — generated by the Planner, fully editable, with per-item version history and snapshot restore
  • Inline editing — edit book metadata, chapter titles/outlines, and add chapters manually without leaving the detail page
  • Version history — chapters, characters, and plot threads all track history with preview and restore; soft-archive instead of delete
  • Archived means archived — an archived chapter, character, or plot thread is kept purely so you can look at it or restore it. It is never sent to a model, never counted in planning or continuity checks, never included in an HTML/FB2/EPUB export or on the public reader page, and agents refuse to write to it
  • Book continuation — create a sequel that copies all settings and inherits ancestor context; RAG and planning reference the full base-book chain
⁠Illustrations
  • AI illustrations with Z-Image or Unsloth Studio — turn on Illustrate this book and the LLM plans illustrations for every finished chapter (what each image shows, where it goes, and its positive and negative prompts, chosen to suit the genre), plus the cover once the book is written. Review and edit the plan, then press Illustrate book to render everything in one image-model run — no LLM calls, so you can unload the LLM first (automatic for Ollama)
  • The same objects, clothes and places in every image — the illustrator keeps a registry of recurring assets (props, outfits, places, creatures, vehicles) and repeats each one's description word for word whenever it appears; characters keep their usual outfit unless the story changes it. Review and edit the registry on the Illustrations page — an edit updates every prompt that shows the asset
  • One style for the whole book — the illustrator writes an art direction once (style, a portrait composition shared by the whole cast, a book-wide negative prompt) from your visual style; it is added word for word to every image and can be edited on the Illustrations page
  • Consistent characters — each character gets a visual description written from their profile and a reference portrait; every image showing them repeats that description, and scenes centred on them reuse the portrait's seed. Re-roll a portrait until the look is right, then re-render their scenes
  • Configurable — image endpoint, steps, CFG, sizes and visual style live in the LLM configuration and presets, so applying a preset switches the text and image models together
⁠Exports & sharing
  • Multiple export formats — HTML (6 colour themes, adjustable font size), EPUB, FB2, and a Metadata document (book info, outlines, planning artifacts, agent messages, token stats). Rendered illustrations and the cover are embedded in all of them, compressed to JPEG to keep the download small
  • Public Library — browse and read published books without logging in (when public mode is enabled); each chapter has its own URL for bookmarking and sharing
⁠Infrastructure
  • Pluggable LLM backend — Ollama (default, local), OpenAI (or any OpenAI-compatible API), or Google AI Studio; settings kept as presets (your default is copied into each new book) and per book
  • Real-time streaming — watch chapters being written token by token via SignalR; planning phases stream with live progressive JSON previews. Only book content streams — a reasoning model's thinking is saved to the chat as one "💭 Thinking" message when each step finishes (even when it fails), and the live view reconnects by itself after a network drop
  • Fail-safe runs — a stalled LLM connection times out instead of hanging the book; double-clicking a start button can't launch two runs; runs interrupted by a server restart are closed out with a note telling you to use Continue
  • Health endpoint — GET /health reports whether the database is reachable (for Docker/uptime checks)
  • Token usage statistics — per-agent prompt and completion token counts, persisted to the database and displayed in a collapsible panel. Calls that error, time out, or are cancelled are recorded too, with a Status column and the failure reason shown next to their partial counts
  • MCP server — built-in Model Context Protocol server at /mcp; connect Claude Desktop, VS Code Copilot, or any MCP client using a per-user API token
  • Multi-user — JWT sign-in (short-lived access token, httpOnly refresh cookie) with an admin role for user management
  • Ollama model management — browse installed models, pull new ones with live progress
  • Global concurrency limit — cap simultaneous agent runs across all books/users with AgentSettings__MaxConcurrentRuns
  • PWA support — installable, works offline for cached content

⁠Architecture

[React SPA — served as static files from ASP.NET wwwroot]
        ↕ REST API + SignalR
[ASP.NET Core 10 API]
        ↕ Direct provider SDKs      ↕ EF Core 10 + Npgsql
[LLM (Ollama / OpenAI / Google)]   [PostgreSQL 16 (pgvector in-DB)]

React is built at image-build time and served from wwwroot/ — there is no separate frontend container at runtime.


⁠Quick Start

⁠Prerequisites
⁠Run with Docker Compose
docker-compose up -d

# App is available at
open http://localhost:5000

On first launch the app shows a Create Admin Account setup screen — the first account registered automatically becomes admin. After signing in, go to Settings to configure your LLM provider and pull an Ollama model.

⁠Run from Docker Hub
docker run -d \
  -p 5000:8080 \
  -e ConnectionStrings__DefaultConnection="Host=<postgres-host>;Port=5432;Database=abook;Username=abook;Password=abook" \
  --add-host host.docker.internal:host-gateway \
  jncchds/abook:latest

PostgreSQL with the pgvector extension must be reachable. The compose file below starts it automatically.

Full docker-compose.yml
services:
  abook-api:
    image: jncchds/abook:latest
    ports:
      - "5000:8080"
    environment:
      - ConnectionStrings__DefaultConnection=Host=postgres;Port=5432;Database=abook;Username=abook;Password=abook
      - ASPNETCORE_ENVIRONMENT=Production
    depends_on:
      postgres:
        condition: service_healthy
    extra_hosts:
      - "host.docker.internal:host-gateway"
    restart: unless-stopped

  postgres:
    image: pgvector/pgvector:pg16
    environment:
      POSTGRES_DB: abook
      POSTGRES_USER: abook
      POSTGRES_PASSWORD: abook
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U abook -d abook"]
      interval: 5s
      timeout: 5s
      retries: 10
    restart: unless-stopped

volumes:
  postgres_data:

⁠Configuration

⁠Environment Variables
VariableDefaultDescription
ConnectionStrings__DefaultConnection—PostgreSQL connection string
ASPNETCORE_ENVIRONMENTDevelopmentProduction disables Swagger
LlmDefaults__ProviderOllamaProvider of the "Default" preset each new user gets (Ollama, OpenAI, GoogleAIStudio); unset = new users start without a preset. The LlmDefaults__* values only seed that preset — books never fall back to them
LlmDefaults__ModelNamellama3Default model name
LlmDefaults__Endpointhttp://host.docker.internal:11434Default LLM endpoint
LlmDefaults__ApiKey—API key (required for OpenAI / GoogleAIStudio; optional for Ollama)
LlmDefaults__EmbeddingModelName—Embedding model for RAG (optional; falls back to chat model)
LlmDefaults__ImageEndpoint—Image server URL for illustrations (e.g. http://gpu-box:8000); applied with LlmDefaults__Provider
LlmDefaults__ImageProviderZImageZImage or UnslothStudio
LlmDefaults__ImageApiKey—Image server API key (Unsloth Studio sk-unsloth-…)
LlmDefaults__ImageModel—Unsloth Studio: image model to load when none is loaded
LlmDefaults__ImageVisualStyle—Default visual style for illustrations
LlmDefaults__ImagesPerChapter2Base number of illustrations per chapter (the illustrator may plan more when a chapter needs them)
Jwt__SigningKeygeneratedKey that signs sign-in tokens — any random string of 32+ characters (e.g. openssl rand -base64 48). If unset, ABook generates one and stores it in the database, logs a ⚠️ warning at startup and shows admins a banner; see Sign-in⁠
Jwt__AccessTokenMinutes15Lifetime of an access token
Jwt__RefreshTokenDays30How long a browser stays signed in without using ABook
AgentSettings__MaxConcurrentRuns3Max simultaneous agent runs across all books/users
PublicModefalseEnable public library (anonymous access to published books)

Changing AgentSettings__MaxConcurrentRuns requires restarting the API process/container.

For local development, copy src/ABook.Api/appsettings.Local.example.json → appsettings.Local.json and fill in your values.

⁠Sign-in

⚠️ Since v0.3.0 ABook uses JWT sign-in instead of cookie sessions. Everyone has to sign in once after updating; books, presets and MCP API tokens are not affected.

The browser gets a short-lived access token (kept in memory, never in storage) and a refresh token in an httpOnly cookie that is only sent to /api/auth, rotated on every use and revocable (sign-out, password change). A browser stays signed in for Jwt__RefreshTokenDays after its last visit.

Tokens are signed with Jwt__SigningKey. You don't have to set it — if it is missing, ABook generates a random key on first start and keeps it in its database, so sessions survive restarts and updates. It does log a ⚠️ warning at every startup and shows admins a banner, because anyone with a copy of the database could then forge sign-ins. To move the key into your configuration:

docker compose exec abook-api dotnet ABook.Api.dll secrets show

Copy the printed Jwt__SigningKey=… line into the environment: section of abook-api, then docker compose up -d (any other random 32+ character value works just as well). Changing the key never signs anyone out: it only invalidates access tokens, and browsers fetch a new one with their refresh cookie.

⁠LLM Providers

LLM settings are kept as presets (the 🔑 Presets page) and per book (each book's Settings page):

ProviderNotes
OllamaDefault. Runs locally; host.docker.internal resolves to the host from inside Docker.
OpenAIProvide an API key and model name (e.g. gpt-4o). Leave endpoint blank for the real OpenAI API; set a custom endpoint for any OpenAI-compatible API (Groq, Together, LM Studio at http://host.docker.internal:1234/v1, etc.).
Google AI StudioNative Gemini connector. Requires an API key from aistudio.google.com⁠. Suggested models: gemini-2.0-flash, gemini-2.5-pro. Embedding model: text-embedding-004.
OpenAI CompatibleFor LM Studio, OpenRouter, vLLM and similar. Reads the stream directly, so non-standard fields such as reasoning_content are captured; reasoning_effort is never sent. API key optional.

How the settings are resolved:

  • Every book has its own LLM settings. When you create a book it gets a copy of your default preset (a sequel copies its base book instead). Later changes to the preset — or picking another default — never affect existing books; change those on the book's Settings page (where Apply Preset copies a preset in).
  • Your default preset is chosen under Settings → Default LLM Preset or with Make default on the Presets page. New users get a private "Default" preset built from the LlmDefaults__* environment variables, if those are set.
  • Shared presets are available to every user (including as their default) but only admins can create, edit or delete them. Admins share a preset with the Shared with everyone toggle; sharing cannot be undone. Anyone who can use a shared preset can see its API keys once it is copied into their books.
  • There is no global or per-user fallback configuration any more (removed in v0.3.0).

Local models often ignore the JSON schema the planning agents send — returning a chapter number as 1.0 or "3", or a text field as a list. The parsers coerce those instead of failing, drop only the entries they genuinely cannot read, and post a ⚠️ note in the book's chat naming each dropped entry and the JSON behind it. If a planning phase does fail outright, the error in the chat quotes what the model actually returned.

⁠Image Generation (Illustrations)

Illustrations are rendered by one of two image providers, chosen under 🎨 Image Generation in a book's LLM settings or in a preset; use Test connection after setting the URL.

  • Z-Image — a self-hosted Z-Image Turbo⁠ server: the FastAPI server in the adate project's compose/zimage folder (POST /generate, GET /images/{id}, GET /health).
  • Unsloth Studio — the image models of an Unsloth Studio⁠ server. Enter the Studio address (with or without /v1) and an API key from Studio's Settings → API. Studio keeps only one model on the GPU, so if it also serves your text model (as an OpenAI Compatible provider at http://<host>:<port>/v1), set Image model too — e.g. unsloth/Z-Image-Turbo-GGUF:z-image-turbo-Q8_0.gguf (a Hub repo id, optionally : plus a checkpoint file). ABook then loads it before rendering whenever no image model is loaded, and Studio unloads the text model to make room. Leave it blank to use whatever is loaded on Studio's Images page.
SettingDefaultNotes
Endpoint—Blank disables rendering (illustrations can still be planned)
Image API key—Unsloth Studio only
Image model—Unsloth Studio only — loaded when Studio has no image model loaded
Visual style—Written into every prompt by the LLM; blank lets it choose one for the genre
Steps / CFG8 / 1.0 (Unsloth Studio: 9 / 0)Z-Image Turbo values; blank on Unsloth Studio lets Studio choose, so set both for other Studio models (e.g. FLUX). Negative prompts only take effect with CFG above 1.0
Portrait / landscape size768×1152 / 1152×768Square images use the portrait area
Timeout300000 msPer image
Unload the LLM before renderingoffOllama only (keep_alive: 0); Unsloth Studio swaps models by itself

Workflow: enable Illustrate this book in the book settings → write the book → open 🎨 Illustrations, review or edit the plan → Render portraits and re-roll any you don't like → Illustrate book.

⁠MCP Access

ABook includes a built-in Model Context Protocol⁠ server at /mcp. Any MCP-compatible client can connect to read and write book content and trigger agent workflows.

Setup:

  1. Open Settings → MCP Access
  2. Generate an API token (or click Regenerate to rotate)
  3. Add the server to your MCP client config using Authorization: Bearer <token>

Claude Desktop (claude_desktop_config.json):

{
  "mcpServers": {
    "abook": {
      "type": "http",
      "url": "http://localhost:5000/mcp",
      "headers": { "Authorization": "Bearer YOUR_TOKEN" }
    }
  }
}

VS Code / GitHub Copilot (.vscode/mcp.json):

{
  "servers": {
    "abook": {
      "type": "http",
      "url": "http://localhost:5000/mcp",
      "headers": { "Authorization": "Bearer YOUR_TOKEN" }
    }
  }
}

⁠Agent Workflow

User creates book (title, premise, genre, target chapters)
         │
         ├─ optional: choose a base book (settings + context copied)
         │
         ▼
  ┌──────────────────────────────────────────────────────┐
  │  Planner — 4-phase pipeline                          │
  │   Phase 1: Story Bible (world-building, tone, rules) │
  │   Phase 2: Character Cards (roles, arcs, goals)      │
  │   Phase 3: Plot Threads (subplots, themes, arcs)     │
  │   Phase 4: Chapter Outlines (title + synopsis each)  │
  │                                                      │
  │   Asks clarifying questions up front; optionally     │
  │   pauses after each phase in Human-assisted mode     │
  └──────────────────────────────────────────────────────┘
         │
         │  ← "Plan Only" stops here so you can review
         │    and edit outlines before clicking "Continue"
         ▼
  For each chapter:
  [Checker]  ── pre-write: checks outline for contradictions
         │
         ▼
  [Writer]  ── writes full chapter prose
         │
         ▼
  [Checker]  ── continuity + style review → structured JSON patches
         │
         ▼
  [Editor]  ── applies patches mechanically (skipped if no issues)
         │
         │  ← optional human pause in assisted mode: read the patched
         │    chapter, edit anything, and steer the rewrite
         ▼
  [Editor]  ── creative rewrite — only when the Checker asked for one
         │      or you gave instructions during the pause
         ▼
      Done ✓

Agents stream tokens via SignalR as they write.

Workflow controls:

ButtonBehaviour
Plan OnlyRuns the Planner and stops so you can review outlines
Write BookFull pipeline from scratch (idempotent — skips completed phases and Done chapters)
ContinueResumes from the first non-Done chapter
Continue PlanningRe-runs only the incomplete planning phases
StopCancels any running agent cleanly

In Human-assisted mode the app also pauses after each planning phase, and once per chapter — after the Editor's mechanical fixes are applied but before the creative rewrite. While it is paused nothing is generating, so you are free to edit the chapter (title, outline and prose), the characters, the plot threads or the story bible; the next step re-reads all of it. Whatever you type in the answer box is passed to the rewrite as author instructions, and the rewrite is skipped altogether when the Checker found nothing to rewrite and you left the box empty.

Chapter edits you make by hand are re-embedded in the background, so retrieval for later chapters sees your text rather than the version the agent wrote.

For books created from a base book, chapter-level RAG retrieval includes embeddings from all ancestor books in the continuation chain.


⁠Development

⁠Prerequisites
⁠Local Setup
# Start PostgreSQL
docker-compose up postgres -d

# Start the UI dev server (proxies API calls to localhost:5000)
cd src/abook-ui
npm install
npm run dev

# Start the API (second terminal)
cd src/ABook.Api
dotnet run --urls http://localhost:5000

The React dev server runs at http://localhost:5173 and proxies /api and /hubs to the ASP.NET server.

⁠Database Migrations
dotnet ef migrations add <MigrationName> --project src/ABook.Infrastructure --startup-project src/ABook.Api
dotnet ef database update --project src/ABook.Infrastructure --startup-project src/ABook.Api

Always use the dotnet ef CLI — never create migration files by hand.

⁠Build Docker Image
docker build -t abook .

The multi-stage Dockerfile builds the React app (Node 20), compiles the .NET API (.NET 10 SDK), and produces a minimal runtime image (ASP.NET 10).

⁠LLM Debug Logging

Set LLM_DEBUG_LOGGING=true to print the full chat history and LLM responses to the application log at Information level.


⁠Tech Stack

LayerTechnology
FrontendReact 19, TypeScript, Vite, Zustand, react-markdown
BackendASP.NET Core 10, C#
LLMOllama, OpenAI SDK, Google AI SDK (per-provider direct calls)
DatabasePostgreSQL 16 via EF Core 10 + Npgsql
Vector storepgvector (in-DB, Pgvector.EntityFrameworkCore)
Real-timeSignalR
AuthJWT (JwtBearer) + rotating refresh tokens, IPasswordHasher<T>, Bearer API token for MCP
MCPModelContextProtocol.AspNetCore 1.2.0 — HTTP/SSE transport, 37 tools
ContainerDocker, Docker Compose

⁠License

MIT⁠

Tag summary

Content type

Image

Digest

sha256:528b243f5…

Size

248.8 MB

Last updated

4 days ago

docker pull jncchds/abook