Why enterprises may want to run Kimi K3 confidentially

Kimi K3 is the verified current model name—not a mistaken reference to Kimi K2. A confidential deployment is an infrastructure pattern for self-hosted weights, not a property of Moonshot’s hosted API.

Abstract botanical artwork combining painted flowers with fine technical linework.

Direct answer

Enterprises may run open-weight Kimi K3 inside an attested CPU-and-GPU confidential environment when prompts, retrieved data, credentials, outputs, or model assets must be hidden from the infrastructure operator. This does not make the model safe or compliant by itself, and Moonshot’s public API should not be described as confidential without a documented TEE and attestation assurance.

Name check: Kimi K3 is real and current

Moonshot AI’s current model catalog identifies kimi-k3 as its flagship model and records the retirement of legacy K2 API models on 25 May 2026. Moonshot also publishes a Kimi K3 repository, technical report, quickstart, and model-specific license. The correct page is therefore Kimi K3, not a correction back to Kimi K2.KIMI-MODELSKIMI-REPOKIMI-REPORT

Why K3 changes the infrastructure conversation

Moonshot describes Kimi K3 as a mixture-of-experts model with 2.8 trillion total parameters, 104 billion activated parameters, and a one-million-token context window. Its official deployment material recommends a supernode-scale configuration with at least 64 accelerators. Those characteristics make confidentiality possible in principle with open weights, but operationally demanding: capacity, interconnect topology, collective communication, startup, and attestation all matter.KIMI-REPORTKIMI-QUICK

Assets worth separating in a K3 threat model
AssetWhy it may be sensitiveWhat a confidential boundary can contribute
Prompts and long contextContracts, research, source code, case files, or internal communications can accumulate in a large contextReduce routine host/operator visibility into runtime plaintext
Retrieved enterprise dataRAG can join multiple high-value systems in one requestGate dataset/session keys on an approved measurement
Tool credentialsAgents may reach databases, ticketing systems, or code repositoriesKeep scoped credentials inside an attested broker; still enforce least privilege
Fine-tuned adapters or derivative weightsThey can encode proprietary behavior and training investmentRelease decryption keys only to approved loaders and device state
Base K3 weightsOpen availability does not remove integrity, provenance, or license obligationsVerify an expected model artifact; confidentiality may be secondary to integrity
Outputs and tracesResults can contain sensitive or derived informationProtect generation in use; separately control recipients, logs, and retention

Three deployment patterns

  1. Customer-controlled confidential cluster. The enterprise operates the verifier and key broker, owns the workload image and policy, and deploys K3 onto verified confidential GPU nodes. This offers the clearest separation from the infrastructure operator and the highest operational burden.
  2. Managed dedicated confidential service. A specialist operates the cluster, while the customer independently verifies evidence and retains control of data/model keys. Contract language should match the technical boundary and support process.
  3. Confidential inference gateway. A smaller attested service protects preprocessing, retrieval, credentials, or prompt transformation while a non-confidential model endpoint receives a minimized payload. This is not end-to-end confidential inference, but it can reduce exposure where full K3-scale CC capacity is unavailable.

A defensible confidential K3 reference architecture

  1. Create a reproducible image containing the exact K3 serving stack, collective libraries, drivers, and policy-controlled configuration.
  2. Launch only on a vendor-supported CPU TEE and GPU confidential-computing combination; verify every accelerator expected by the distributed job.
  3. Validate fresh CPU and GPU evidence, security versions, debug state, firmware, and workload measurement through an owned or trusted verifier.
  4. After successful appraisal, issue short-lived keys for encrypted model artifacts, request sessions, and narrowly scoped tools.
  5. Keep prompts, KV cache, intermediate state, and output generation inside the approved device/VM path; redact operational telemetry.
  6. Revoke or withhold new keys when measurements drift, a node is replaced, an accelerator drops from the attested set, or a security baseline changes.
NVIDIA-DEPLOYNVIDIA-ATTEST

Limitations and procurement caveats

  • Scale: K3’s recommended multi-accelerator footprint narrows the pool of supported confidential infrastructure and raises the importance of topology verification.
  • Performance: encrypted CPU–GPU transfers, protected I/O, attestation startup, and disabled optimization paths can add overhead. Benchmark the exact serving engine and context profile.
  • Software security: an attested vulnerable inference server is still vulnerable. Model and tool-layer attacks remain in scope for application controls.
  • Multi-node boundary: attesting one VM is not enough when a request is sharded across many nodes and accelerators. Admission, membership, and keying must cover the distributed job.
  • Availability: an infrastructure operator may still stop, starve, or reset the service.
  • Licensing: Kimi K3 uses the Kimi K3 License. Legal teams should review the official terms for the intended distribution and commercial model rather than assuming K2’s Modified MIT terms apply.
KIMI-LICENSENVIDIA-DEPLOY

When confidential K3 is justified

The strongest case combines valuable private context, a need to use third-party infrastructure, and the ability to operate meaningful verification and key-release policy. If the organization controls the physical cluster and administrators already sit inside the trust boundary, encrypted storage, strict access control, and supply-chain hardening may address more immediate risks. If model scale makes full confidential deployment impractical, reduce or partition the sensitive context before treating “confidential K3” as a procurement requirement.

Frequently asked questions

Is Kimi K3 an official model name?

Yes. As of 22 August 2026, Moonshot AI’s official model catalog lists kimi-k3 as its flagship model, and Moonshot publishes an official repository, technical report, quickstart, and license.

Is the hosted Kimi API confidential computing?

Moonshot does not make that claim in its public documentation. A hosted API needs an explicit TEE boundary, verifiable evidence, and an attestation-bound key-release design before it can be described as confidential computing.

Do open weights remove the need for confidentiality?

Not necessarily. Prompt data, enterprise retrieval, tool credentials, fine-tuned adapters, intermediate state, and outputs may remain sensitive. Open availability can make base-weight secrecy less important while integrity and provenance still matter.

Sources

  1. KIMI-MODELS
    ModelsMoonshot AI / Kimi Platform — Current catalog and K2 retirement notice
  2. KIMI-REPO
  3. KIMI-REPORT
  4. KIMI-QUICK
    Kimi K3 QuickstartMoonshot AI / Kimi Platform
  5. KIMI-LICENSE
    Kimi K3 LicenseMoonshot AI
  6. NVIDIA-DEPLOY
  7. NVIDIA-ATTEST

Relevant GPU availability

Verified specifications, confidential-mode support, and public listings for the accelerators this post covers.

Ready to reserve capacity?

Confidential Nodes matches bare-metal confidential GPU nodes to workloads, with verified provider data behind every listing.