Skip to content

BUG: K8s provider cert mismatch (x509) when tofu plan/apply uses SSH tunnel #7

Description

@jgruberf5

Summary

When the OpenTofu K8s tunnel (tofu_k8s_tunnel.py) opens an SSH tunnel for tofu plan/tofu apply, it writes a provider_override.tf that points the kubernetes provider at https://127.0.0.1:{tunnel_port}. However, the EKS API server certificate is only valid for the real endpoint hostname (e.g. *.eks.amazonaws.com) and internal K8s service IPs — not 127.0.0.1. This causes every kubernetes_* resource refresh to fail with:

x509: certificate is valid for 172.20.0.1, 10.0.20.202, 172.16.45.87, not 127.0.0.1

This blocks all plan/apply operations on modules that use the kubernetes provider through SSH tunnels.

Root Cause

In backend/services/tofu_k8s_tunnel.py, the _write_provider_override() function writes:

provider "kubernetes" {
  host                   = "https://127.0.0.1:{local_port}"
  cluster_ca_certificate = base64decode(...)
  token                  = ...
}

The host is 127.0.0.1 but TLS verification checks the certificate's SANs against the connection hostname, which fails.

Note: The Forge API/UI K8s client (cluster_utils.py, kubernetes/_base.py) handles this correctly by setting insecure-skip-tls-verify on the kubeconfig. But the Terraform provider override does not have an equivalent.

Fix

Add tls_server_name to the provider override. The Terraform kubernetes provider supports this attribute — it sets the TLS SNI header so cert verification checks against the original hostname instead of 127.0.0.1:

def _write_provider_override(work_dir, local_port, eks_endpoint):
    from urllib.parse import urlparse
    _parsed_ep = urlparse(eks_endpoint)
    tls_server_name = _parsed_ep.hostname or ""

    override_content = f'''
provider "kubernetes" {{
  host                   = "https://127.0.0.1:{local_port}"
  tls_server_name        = "{tls_server_name}"
  cluster_ca_certificate = base64decode(data.aws_eks_cluster.cluster.certificate_authority[0].data)
  token                  = data.aws_eks_cluster_auth.cluster.token
}}
'''

The eks_endpoint is already passed to the function (currently only used in a comment). This parses the hostname and uses it for TLS SNI.

Status

  • Fix applied on server (10.145.33.194) in ~/git/bnk-forge/backend/services/tofu_k8s_tunnel.py, containers rebuilt — plan/apply works through SSH tunnel
  • Not committed to git — needs PR
  • Should also consider adding tls_server_name for the helm provider if/when helm provider overrides are added

Impact

Without this fix, any module with kubernetes provider resources (e.g. high-performance-nodes with multus/sriov/dpdk k8s manifests) cannot plan or apply when the cluster is accessed via SSH tunnel.


Migrated from sp-prod-field/bnk-forge #22 (opened 2026-04-14; original labels: bug). That repository is archived and read-only.
Bare #NNN references in the text above refer to issues and PRs in the original repository, not to numbering here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaves incorrectlykubernetesK8s engine, Helm, CRDs, cluster operationsopentofuOpenTofu/Terraform engine and modulesseverity:highBlocks a core workflow, needs manual intervention, or is a security gap

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions