Kubernetes Deployment
Deploy the Xberg REST API server (xberg serve) on Kubernetes with the official Helm chart. The chart is a single-service deployment — one stateless workload plus an optional cache, ingress, and autoscaler. For a managed, multi-tenant platform with a work queue, observability, and billing, see Xberg Enterprise.
Install
Section titled “Install”The chart is published as an OCI artifact to GitHub Container Registry:
helm install xberg oci://ghcr.io/xberg-io/charts/xberg --version 1.0.0-rc.25This runs the full image (ghcr.io/xberg-io/xberg) in API-server mode on port 8000, exposed through a ClusterIP Service on port 80.
The chart is also listed on Artifact Hub, where you can browse every version, its values, and the generated values schema.
Verify the chart
Section titled “Verify the chart”Every published chart is signed with cosign using keyless signing (Sigstore OIDC + the Rekor transparency log). Verify a release before installing:
cosign verify \ ghcr.io/xberg-io/charts/xberg:1.0.0-rc.25 \ --certificate-identity-regexp '^https://github.com/xberg-io/xberg/.github/workflows/publish-helm.yaml@.*$' \ --certificate-oidc-issuer https://token.actions.githubusercontent.comThe identity certificate ties each signature to the publish-helm.yaml workflow in this repository, so a successful verification proves the chart was built and pushed by Xberg’s release pipeline.
Configure
Section titled “Configure”Override defaults with a values.yaml file:
image: # Empty tag defaults to the chart appVersion. Use "core" for the minimal # image (no pre-downloaded models) or "latest" for the full image. tag: ""
xberg: logLevel: "info" ocrLanguage: "eng"
resources: requests: memory: "1Gi" cpu: "1000m" limits: memory: "4Gi" cpu: "2000m"
ingress: enabled: true className: "nginx" hosts: - host: xberg.example.com paths: - path: / pathType: Prefix tls: - secretName: xberg-tls hosts: - xberg.example.com
autoscaling: enabled: true minReplicas: 1 maxReplicas: 10 targetCPUUtilizationPercentage: 80helm install xberg oci://ghcr.io/xberg-io/charts/xberg \ --version 1.0.0-rc.25 \ -f values.yamlCache and replicas
Section titled “Cache and replicas”Embedding and OCR models range from ~90 MB to 1.2 GB and are re-downloaded on every pod restart without a cache. The chart enables a ReadWriteOnce PVC (cache.enabled: true) mounted at /app/.xberg (with HF_HOME under it) to persist them.
Upgrade and uninstall
Section titled “Upgrade and uninstall”helm upgrade xberg oci://ghcr.io/xberg-io/charts/xberg --version 1.0.0-rc.25 -f values.yamlhelm uninstall xbergThe cache PVC carries helm.sh/resource-policy: keep, so it survives an uninstall — delete it manually if you no longer need the cached models.
What’s included
Section titled “What’s included”| Resource | Description | Conditional |
|---|---|---|
| Deployment | API server with health probes and a hardened, non-root, read-only-root security context | Always |
| Service | ClusterIP on port 80 → container 8000 | Always |
| ServiceAccount | Dedicated service account | serviceAccount.create |
| PersistentVolumeClaim | Cache for models and downloaded assets | cache.enabled |
| Ingress | HTTP(S) ingress with optional TLS | ingress.enabled |
| HorizontalPodAutoscaler | CPU/memory-based autoscaling | autoscaling.enabled |
| PodDisruptionBudget | Availability during voluntary disruptions | podDisruptionBudget.enabled |
All values are documented in the chart’s values.yaml and validated on install against the bundled values.schema.json, so a malformed override fails fast with a clear error.
Next steps
Section titled “Next steps”- Docker Deployment — image variants and execution modes
- API Server — endpoint reference
- OCR — backends and language configuration