How It Works
Private AI, Built on Your Infrastructure
KoolBox is a complete AI platform that runs entirely within your environment. Connect your GPU nodes, deploy models, build knowledge bases, and automate workflows — all without sending data outside your organization.

KoolBox Control Panel — manage nodes, models, knowledge bases, and AI services
01
Knowledge Bases
Upload documents (PDF, DOCX, TXT, HTML), fetch URLs, or connect APIs. KoolBox parses, chunks, and indexes your content into searchable knowledge bases. Employees get instant answers grounded in your approved company information.
- Supports PDF, DOCX, TXT, MD, HTML, CSV, JSON
- Image processing with vision models
- Automatic chunking with configurable overlap
- Deploy to remote GPU nodes for local inference

02
RAG Pipeline Builder
Design custom retrieval-augmented generation pipelines with a visual drag-and-drop editor. Connect document loaders, text splitters, embeddings, vector stores, retrievers, and LLM inference nodes into powerful processing chains.
- 14+ node types: loaders, splitters, embeddings, retrievers
- Guardrails and PII redaction built-in
- Conditional routing and custom scripts
- Runs on your own GPU nodes via SSH

03
Workflow Automation
Automate multi-step business processes with AI. Build workflows that combine AI tasks, human-in-the-loop approvals, external integrations, and conditional branching — all executing on your infrastructure.
- AI Task nodes run inference on remote GPUs
- Integration nodes: Slack, REST, Stripe, and more
- Form nodes for human-in-the-loop steps
- Scheduled jobs: cron, hourly, daily, weekly

04
AI Model Deployment
Deploy open-source AI models to your GPU nodes with one click. KoolBox manages the full lifecycle — pulling from Ollama registry or Hugging Face, deploying via SSH, and monitoring status. Models run entirely on your hardware.
- Ollama registry: Llama, Mistral, Phi, and more
- Hugging Face GGUF model support
- Auto-detect GPU, VRAM, and CUDA version
- No data leaves your infrastructure

05
Infrastructure Management
Connect GPU servers via SSH and manage them from a single control panel. KoolBox auto-discovers hardware capabilities and provides real-time monitoring of your AI infrastructure.
- SSH-based secure node connections
- Auto-detect CPU, RAM, GPU, VRAM, CUDA
- Deploy models and KBs to specific nodes
- Health monitoring and service status

06
AI Chat & Assistants
Deploy scoped AI chatbots for different teams or use cases. Each chatbot has its own system prompt, model, knowledge base, and access controls. Every answer is grounded in your approved data.
- Per-team chatbots with custom system prompts
- Knowledge base integration for grounded answers
- File attachments (documents and images)
- API keys for programmatic access

Architecture
How It All Connects
The KoolBox control panel runs locally (or in your private cloud). AI inference runs on remote GPU nodes connected via SSH. Your data never leaves your infrastructure.
Next.js app + PostgreSQL. Manages users, pipelines, chatbots, and orchestration.
Secure communication channel. No ports exposed, no data in transit outside your network.
Ollama runs model inference. Models and KB data stored locally on the node.
Ready to Deploy Private AI?
Start with the free Community plan — 1 node, full AI chat, no credit card required.
FAQ