Local Second Brain

Second Brain Project — Daily Report

Date: August 22, 2026

Executive Summary

Established the initial working architecture for a fully local AI Second Brain running on an Ubuntu/NVIDIA workstation. The system uses Qwen3 8B as the primary LLM, Qwen3-Embedding 0.6B for semantic embeddings, Qdrant as the vector database, and Open WebUI as the browser-based user interface.

The system successfully performs end-to-end Retrieval-Augmented Generation (RAG) against documents stored on NAS. A PDF was ingested, converted to embeddings, stored in Qdrant, semantically retrieved, and successfully queried with Qwen.

The ingestion pipeline was subsequently improved to preserve PDF page boundaries and page metadata, providing a foundation for page-level source citations and more accurate document retrieval.

Open WebUI was connected to the external Qdrant knowledge collection. Retrieval through Open WebUI is functional, although answer quality and external-knowledge integration still require tuning. Direct queries using the custom rag.py pipeline currently provide more predictable results.

The project now has a functional local Second Brain foundation with the original documents remaining on NAS and the rebuildable AI indexes maintained separately on the AI server.

Hardware

AI Server

  • OS: Ubuntu Linux
  • GPU: NVIDIA GeForce RTX 3070 Ti
  • GPU VRAM: 8 GB
  • NVIDIA Driver: 595.84
  • CUDA: 13.2
  • System RAM: ~60 GiB usable
  • Swap: 8 GiB
  • Local system storage: 218 GB
  • Local free space at project start: ~191 GB

Qwen3 8B was verified with ollama ps to run 100% on the GPU without CPU offload.

Initial Ollama context size was 4096 tokens. Work began to increase this to 8192 for improved RAG performance.

Storage Architecture

Original Second Brain documents are maintained on NAS:

/mnt/nas/project/secondbrain/
├── inbox/
├── documents/
├── projects/
├── reference/
└── archive/

AI infrastructure is kept separately on the Ubuntu server:

/opt/second-brain/
├── qdrant/
└── ingestion/
    ├── venv/
    ├── ingest.py
    ├── rag.py
    └── state.db

This intentionally separates the authoritative knowledge repository from disposable/rebuildable AI indexes.

Models

Primary LLM

qwen3:8b

Running through Ollama and verified at 100% GPU utilization.

Embedding Model

qwen3-embedding:0.6b

Produces 1024-dimensional vectors used by Qdrant for semantic retrieval.

Software / Tools

  • Ubuntu Linux
  • NVIDIA CUDA / NVIDIA driver
  • Ollama
  • Qwen3 8B
  • Qwen3-Embedding 0.6B
  • Docker
  • Open WebUI
  • Qdrant
  • Python 3 virtual environment
  • SQLite
  • qdrant-client
  • requests
  • pypdf
  • python-docx
  • python-pptx
  • BeautifulSoup4
  • lxml
  • systemd / systemd timers

Work Completed

  1. Verified NVIDIA GPU configuration using nvidia-smi.

  2. Confirmed RTX 3070 Ti with 8 GB VRAM, CUDA 13.2, and approximately 60 GiB system RAM.

  3. Installed and configured Ollama.

  4. Downloaded and tested:

qwen3:8b
  1. Verified with ollama ps that Qwen3 8B runs:
PROCESSOR: 100% GPU
  1. Installed Open WebUI in Docker.

  2. Resolved Docker permissions for the rw user.

  3. Configured Ollama to listen on an interface accessible from Docker.

  4. Connected Open WebUI to the host Ollama service.

  5. Verified qwen3:8b appears and operates through Open WebUI.

  6. Selected the existing NAS location as the authoritative Second Brain document repository:

/mnt/nas/project/secondbrain
  1. Created the initial Second Brain NAS directory organization.

  2. Deployed Qdrant in Docker with persistent local storage under:

/opt/second-brain/qdrant
  1. Installed:
qwen3-embedding:0.6b
  1. Verified Ollama's embedding API.

  2. Created a Python virtual environment for the ingestion system.

  3. Developed the first document ingestion script.

  4. Implemented:

  • Recursive document discovery
  • Text extraction
  • Document chunking
  • Qwen embedding generation
  • Qdrant vector insertion
  • Source metadata
  • Deterministic Qdrant point IDs
  1. Created a test knowledge document and successfully inserted it into Qdrant.

  2. Created search.py to directly perform semantic searches against Qdrant.

  3. Verified semantic retrieval using questions that did not exactly match the source wording.

  4. Created rag.py to implement the complete RAG workflow:

Question
   ↓
Qwen3-Embedding
   ↓
Qdrant semantic search
   ↓
Top relevant chunks
   ↓
Qwen3 8B
   ↓
Answer
  1. Verified RAG answers and tested behavior when information was absent from the knowledge base.

  2. Expanded document ingestion support to:

  • TXT
  • Markdown
  • PDF
  • DOCX
  • PPTX
  • HTML
  1. Installed supporting parsing packages:
  • pypdf 6.16.1
  • python-docx 1.2.0
  • python-pptx 1.0.2
  • beautifulsoup4 4.15.0
  • lxml 6.1.2
  1. Added SHA-256 document hashing and SQLite state tracking.

  2. Added detection of:

  • New documents
  • Modified documents
  • Unchanged documents
  • Deleted documents
  1. Successfully ingested a real solar contract PDF from the NAS.

  2. Successfully queried information from that PDF using the custom RAG pipeline.

  3. Began configuring automatic ingestion using a systemd oneshot service and timer.

  4. Identified the need for NAS mount protection so a temporary NAS outage cannot be mistaken for deletion of all Second Brain documents.

  5. Connected Open WebUI to the existing external Qdrant secondbrain collection.

  6. Configured Open WebUI to use the same embedding model as the custom ingestion system:

Engine: Ollama
Model: qwen3-embedding:0.6b
  1. Diagnosed external Qdrant payload mappings in Open WebUI.

  2. Verified directly through the Qdrant API that PDF chunks and metadata are correctly stored.

  3. Determined that Open WebUI retrieval is functioning, but answer quality currently trails the custom rag.py implementation.

  4. Modified PDF ingestion to preserve individual PDF pages instead of creating chunks spanning page boundaries.

  5. Added Qdrant metadata including:

filename
source
relative_path
page
page_chunk
chunk
section_type
modified
  1. Re-indexed the PDF using the new page-aware ingestion system.

  2. Verified Qdrant now contains page-specific chunks, for example:

filename: -*Redacted*-.pdf
page: 3
page_chunk: 0
section_type: page

Current Architecture

             NAS
              │
     Original Documents
              │
              ▼
        Python Ingester
              │
       ┌──────┴───────┐
       │              │
 Text Extraction   Metadata
       │
       ▼
 Page-aware Chunking
       │
       ▼
Qwen3-Embedding 0.6B
       │
  1024-d vectors
       │
       ▼
     Qdrant
       │
       ▼
 Semantic Retrieval
       │
       ▼
    Qwen3 8B
       │
       ▼
      Answer

Open WebUI provides the browser-based interface and is being integrated with this external knowledge system.

Current Status

Working:

  • Local Qwen inference
  • NVIDIA GPU acceleration
  • Open WebUI
  • NAS document storage
  • Qdrant vector database
  • Qwen embeddings
  • TXT/Markdown ingestion
  • PDF ingestion
  • DOCX/PPTX/HTML support
  • Semantic search
  • RAG
  • File change tracking
  • Page-aware PDF chunks
  • External Qdrant connection from Open WebUI

Needs further work:

  • Improve Open WebUI RAG answer quality
  • Increase/tune Qwen context window
  • Finalize automatic ingestion
  • Add NAS mount-loss protection
  • Improve PDF table/layout extraction
  • Improve citations in generated answers
  • Add OCR for scanned PDFs
  • Add image/screenshot ingestion

Recommended Next Steps

  1. Update rag.py to expose PDF page numbers directly to Qwen and produce filename/page citations.
  2. Complete and validate the systemd automatic ingestion timer.
  3. Add NAS availability protection before automatic synchronization modifies Qdrant.
  4. Tune Open WebUI Top-K retrieval and RAG prompting.
  5. Evaluate a layout-aware PDF parser such as Docling for contracts, tables, invoices, and structured documents.
  6. Add OCR fallback for scanned/image-only PDFs.
  7. Benchmark larger Qwen models using the server's 60 GiB RAM for partial CPU/GPU offload.
  8. Eventually add multimodal ingestion for photographs, diagrams, screenshots, and other visual material.

End-of-day status: A functional, fully local Second Brain RAG system is operational and successfully answering questions from real documents stored on NAS.