Virtual Machines
KVM, OpenStack, Proxmox VE and KubeVirt
GPU Infrastructure
We build GPU infrastructure for AI inference in your environment: GPU servers and VMs with passed-through GPUs, local language models, embeddings and retrieval. Your data never leaves your premises.
Features
Dedicated GPUs on bare metal or through PCIe passthrough in virtual machines.
Open language models behind an OpenAI-compatible API, with vLLM.
Vector database and embedding models for RAG, right next to inference.
Prompts, documents and results never leave your infrastructure.
Models and datasets on block, file or object storage.
GPU utilisation, memory and throughput at a glance.
Use cases
Language models for knowledge, documents and support, without external AI services.
Classification, extraction and semantic search across large document collections.
Test and compare models before they go to production.
Services
KVM, OpenStack, Proxmox VE and KubeVirt
Automated provisioning of physical servers
S3-compatible storage on Ceph
An architecture review, a new environment from the ground up or support in operations: talk directly to the engineers who will deliver it.