Private AI, Built on Your Infrastructure

KoolBox is a complete AI platform that runs entirely within your environment. Connect your GPU nodes, deploy models, build knowledge bases, and automate workflows — all without sending data outside your organization.

KoolBox dashboard showing AI nodes, services, and system health

KoolBox Control Panel — manage nodes, models, knowledge bases, and AI services

Knowledge Bases

Upload documents (PDF, DOCX, TXT, HTML), fetch URLs, or connect APIs. KoolBox parses, chunks, and indexes your content into searchable knowledge bases. Employees get instant answers grounded in your approved company information.

  • Supports PDF, DOCX, TXT, MD, HTML, CSV, JSON
  • Image processing with vision models
  • Automatic chunking with configurable overlap
  • Deploy to remote GPU nodes for local inference
KoolBox knowledge base management interface

RAG Pipeline Builder

Design custom retrieval-augmented generation pipelines with a visual drag-and-drop editor. Connect document loaders, text splitters, embeddings, vector stores, retrievers, and LLM inference nodes into powerful processing chains.

  • 14+ node types: loaders, splitters, embeddings, retrievers
  • Guardrails and PII redaction built-in
  • Conditional routing and custom scripts
  • Runs on your own GPU nodes via SSH
KoolBox RAG pipeline visual builder

Workflow Automation

Automate multi-step business processes with AI. Build workflows that combine AI tasks, human-in-the-loop approvals, external integrations, and conditional branching — all executing on your infrastructure.

  • AI Task nodes run inference on remote GPUs
  • Integration nodes: Slack, REST, Stripe, and more
  • Form nodes for human-in-the-loop steps
  • Scheduled jobs: cron, hourly, daily, weekly
KoolBox workflow automation builder

AI Model Deployment

Deploy open-source AI models to your GPU nodes with one click. KoolBox manages the full lifecycle — pulling from Ollama registry or Hugging Face, deploying via SSH, and monitoring status. Models run entirely on your hardware.

  • Ollama registry: Llama, Mistral, Phi, and more
  • Hugging Face GGUF model support
  • Auto-detect GPU, VRAM, and CUDA version
  • No data leaves your infrastructure
KoolBox AI model deployment interface

Infrastructure Management

Connect GPU servers via SSH and manage them from a single control panel. KoolBox auto-discovers hardware capabilities and provides real-time monitoring of your AI infrastructure.

  • SSH-based secure node connections
  • Auto-detect CPU, RAM, GPU, VRAM, CUDA
  • Deploy models and KBs to specific nodes
  • Health monitoring and service status
KoolBox server and node management

AI Chat & Assistants

Deploy scoped AI chatbots for different teams or use cases. Each chatbot has its own system prompt, model, knowledge base, and access controls. Every answer is grounded in your approved data.

  • Per-team chatbots with custom system prompts
  • Knowledge base integration for grounded answers
  • File attachments (documents and images)
  • API keys for programmatic access
KoolBox AI chat and assistant configuration

How It All Connects

The KoolBox control panel runs locally (or in your private cloud). AI inference runs on remote GPU nodes connected via SSH. Your data never leaves your infrastructure.

Control Panel

Next.js app + PostgreSQL. Manages users, pipelines, chatbots, and orchestration.

SSH

Secure communication channel. No ports exposed, no data in transit outside your network.

GPU Nodes

Ollama runs model inference. Models and KB data stored locally on the node.

Ready to Deploy Private AI?

Start with the free Community plan — 1 node, full AI chat, no credit card required.

Frequently Asked Questions