A modular, local-first AI agent system combining MCP (Model Context Protocol), RAG (Retrieval-Augmented Generation), and Skills (execution layer) into a unified architecture. Runs entirely on your machine with no cloud dependency.
Optimized for NVIDIA RTX 3050 (4GB VRAM), 16GB RAM, 512GB SSD, Intel CPU.
┌─────────────────────────────────────────────────────┐
│ User Interface │
│ Streamlit / Gradio / FastAPI + React │
├─────────────────────────────────────────────────────┤
│ Orchestrator │
│ Routes queries → decides → executes + responds │
├──────────┬──────────────────┬───────────────────────┤
│ RAG │ MCP │ Skills │
│ Layer │ Connection │ Execution │
│ │ Layer │ Layer │
├──────────┴──────────────────┴───────────────────────┤
│ LLM (Ollama) │
│ Llama 3.1 8B / Mistral 7B / Qwen 2.5 7B │
│ RTX 3050 GPU │
└─────────────────────────────────────────────────────┘
- Ollama for local model serving
- Models: Llama 3.1 8B, Mistral 7B, Qwen 2.5 7B
- CUDA-accelerated inference on RTX 3050
- ChromaDB as local vector store
- LangChain / LlamaIndex for orchestration
- Embeddings:
all-MiniLM-L6-v2orbge-base-en-v1.5 - Document ingestion pipeline (PDF, DOCX, TXT, HTML)
- File system MCP server
- SQLite MCP server
- Custom industrial MCP servers (OPC UA, Modbus)
- Git MCP server for versioning
- Python code execution (sandboxed)
- File read/write operations
- Data analysis (pandas, numpy, matplotlib)
- Report generation (Word, Excel, PDF)
- Database queries (SQLite, PostgreSQL)
- OPC UA / Modbus data acquisition
- Streamlit (rapid prototyping)
- Gradio (quick demos)
- Or FastAPI + React (production-grade)
# 1. Install Ollama
# Download from https://ollama.com
# 2. Pull a model
ollama pull llama3.1:8b
# 3. Create Python virtual environment
python -m venv .venv
.venv\Scripts\activate
# 4. Install dependencies
pip install -r requirements.txt
# 5. Run the application
streamlit run src/main.pymcp-rag-skills-agent/
├── config/ # Configuration files
│ ├── settings.yaml # Main configuration
│ └── skills_registry.yaml # Skills catalog
├── data/
│ ├── documents/ # Source documents for RAG
│ ├── chroma_db/ # Vector database storage
│ └── logs/ # Application logs
├── src/
│ ├── llm/ # LLM integration
│ │ └── ollama_client.py
│ ├── rag/ # RAG pipeline
│ │ ├── embeddings.py
│ │ ├── vector_store.py
│ │ └── retriever.py
│ ├── mcp_servers/ # MCP protocol servers
│ │ ├── file_server.py
│ │ ├── sqlite_server.py
│ │ └── mcp_client.py
│ ├── skills/ # Execution skills
│ │ ├── registry.py
│ │ ├── python_executor.py
│ │ └── data_analyzer.py
│ ├── interface/ # UI layer
│ │ ├── app.py # Streamlit app
│ │ └── api.py # FastAPI app
│ └── main.py # Entry point
├── tests/
├── docs/
├── requirements.txt
├── README.md
├── ARCHITECTURE.md
└── CAHIER_DES_CHARGES.md
Edit config/settings.yaml:
llm:
provider: ollama
model: llama3.1:8b
temperature: 0.3
max_tokens: 2048
rag:
embedding_model: all-MiniLM-L6-v2
chunk_size: 512
chunk_overlap: 64
top_k: 5
mcp:
servers:
- file_system
- sqlite
- git
skills:
sandbox: true
timeout: 30- RAG on equipment manuals, procedures, standards
- Skills: OEE calculations, report generation, SQL queries
- MCP: local file access, database connectivity
- RAG on production history
- Skills: Python analysis (pandas), visualizations
- MCP: database connections, file I/O
- RAG on templates and existing docs
- Skills: Word/Excel/PDF creation
- MCP: Git integration for versioning
MIT