Skip to content

[New Project Proposal]: - Thalamus #3

Description

@auhlig

Contact Details

arno.uhlig@sap.com

Mission Statement

The goal of the project is to create sovereign and vendor-neutral AI infrastructure for LLM deployments.
I'm requesting the creation of a thalamus-ai GitHub organization.

Project Description

Thalamus

Thalamus is a vendor-neutral, Kubernetes-native inference service for sovereign LLM deployments. Weights, prompts, and context are protected and stay in the deployment perimeter.

Thalamus is built on llm-d, vLLM, the Gateway API inference extension, and Cortex.
Supports Nvidia, AMD, and Intel infrastructure.

Value Propositions

  • Regulatory unblock. Serves regulated and public-sector workloads that cannot use hyperscaler AI providers.
  • No vendor lock-in. No dependency on a single hyperscaler, LLM vendor, or GPU infrastructure provider. Swap models and hardware behind a stable API.
  • IP and data protection. Confidential computing keeps model weights and customer context inside attested GPU boundaries.
  • One stack, many targets. The same stack runs in datacenters, sovereign cloud, and air-gapped satellite sites.

Features

  • OpenAI-compatible inference API with streaming
  • Declarative model management via a Kubernetes Custom Resources including full lifecycle management
  • Model-aware routing through the Gateway API Inference Extension
  • KV-cache aware load balancing via the llm-d Endpoint Picker
  • Autoscaling driven by time-to-first-token and queue-depth metrics
  • Multi-vendor GPU support (Nvidia, AMD, Intel)
  • Observability via Prometheus, OpenTelemetry, and GPU exporters. Integrates natively with Greenhouse.
  • Ready for air-gapped satellite deployments

Alignment to NeoNephos mission

Thalamus is a vendor-neutral, Kubernetes-native inference service for sovereign LLM deployments.

First, sovereignty: Thalamus removes the dependency on hyperscaler AI services and supports air-gapped and regulated deployments. This is the lock-in problem, the NeoNephos sovles.

Second, openness and interoperability: Apache-2.0, built on open upstream components (llm-d, vLLM, Gateway API Inference Extension), OpenAI-compatible API, multi-vendor GPU support (NVIDIA, AMD, Intel), runs on any Kubernetes, leverages and integrates with existing ApeiroRA components.

Third, ApeiroRA: Thalamus complements the existing ApeiroRA and foundation scope by providing the missing sovereign AI inference layer to the stack.

Benefit to NeoNephos

Thalamus addresses a critical gap in the NeoNephos Foundation by providing sovereign GenAI inference.
It arrives production-grade, Apache-2.0, ApeiroRA-integrated (CobaltCore, Greenhouse, OCM, Gardener).
It pulls the European inference community (llm-d, vLLM, GIE operators) into the foundation and gives sponsors a concrete AI workload story beyond infrastructure.

The risk of not accepting this that the AI layer of the European sovereign stack gets defined elsewhere, e.g. by a single vendor, vendor consortium, or a hyperscaler-aligned project. ApeiroRA ships without an open inference reference and members buy proprietary stacks instead.

Benefit to Project

Thalamus is already part of NeoNephos today as a CobaltCore sub-component.
While it was incubated there, the scope doesn't fit into CobaltCore.
A standalone status ensures visibility of the LLM inference service, the technical direction instead of being nested in CobaltCore, enables direct access to sponsors, overnment channels, and collaboration with partners.

Is this a new project or an existing one?

Existing Project

Project Leadership

Arno Uhlig - TSC Chairperson
Valentin Küchler - TSC Member
Henry Richter - TSC Member

Project Sponsors

SAP, NeoNephos as Thalamus is being incubated in CobaltCore.

Release Methodology

We use semantic versioning.

Current Infrastructure

Security Response

GitHub vulnerability reporting is enabled, tracked via GitHub issues, and responded to timely based on severity.

External Dependencies

See golang sbom

Infrastructure Needs

Workers for GitHub actions, GitHub container registry for OCI images

Current Project License

Apache 2.0

Project Type

Code Project

Code of Conduct

https://github.com/cobaltcore-dev/thalamus#code-of-conduct

End Users

SAP Cloud Infrastructure, AMD

Trademark Inventory

Thalamus, thalamus-ai

TAC Supporters

Florian Müller, SAP

Good Faith

  • I agree

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions