Kubernetes API Server Adventures: Part 1 — Extending The Kubernetes API
Introduction
Something that I love about Kubernetes is how unapologetically “API-first” it is. Everything is an API call, everything is declarative, and the whole control plane is designed less like a monolithic piece of software and more like a framework you can extend.
Now, it also gives you two main ways to extend it. You can take the lightweight route with CustomResourceDefinitions (CRDs): define your new resource types in YAML, hand them to the API server, and suddenly your cluster speaks a new dialect without a single line of Go code. Or, if you need to go beyond what etcd and schemas can offer, you can reach for the heavyweight option: an extension API server. That’s when you’re effectively writing your own Kubernetes-style API server, plugging it into the aggregation layer, and letting the main kube-apiserver front your service just like it does for built-in groups.
Both approaches are first-class citizens. But they carry very different trade-offs — in complexity, in control, and in operational cost. We’ll see how to extend the Kubernetes API, from CRDs to extension API servers, and which path makes sense for any given scenario.

Disclaimer: I like making mistakes :). If you’ve noticed one, please let me know and correct me.
Custom Resource Definitions (CRDs)
These are the bread-and-butter of Kubernetes extensibility:
- Declarative — you create a CRD YAML, and the API server instantly serves new endpoints (e.g.,
/apis/mygroup/v1/myfoos). No extra coding needed. - You get all of Kubernetes’ built-in features: stored in etcd, RBAC, validation (via OpenAPI schemas), and
kubectlsupport.
If you’ve spent any time building “operators,” you probably already know CRDs are the workhorse of Kubernetes extensibility. What still surprises me is how much the API server itself does for you once you define a CRD well. You’re not just bolting on a JSON blob; you’re teaching the control plane a new resource type with discovery, OpenAPI, RBAC, watch semantics, strategic merge/SSA behavior, table output, and even schema-aware validation — all without patching kube-apiserver.
The core idea is that you declare a CustomResourceDefinition in the apiextensions.k8s.io/v1 API, provide a structural OpenAPI v3 schema, and the main API server starts serving and storing your resources like first-class citizens. Here’s how a CRD looks like:
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
metadata:
name: widgets.example.com
spec:
group: example.com
scope: Namespaced
names:
kind: Widget
plural: widgets
singular: widget
shortNames:
- wdg
versions:
- name: v1
served: true
storage: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
size:
type: string
enum: ["S","M","L"]
default: "M"
replicas:
type: integer
minimum: 0
default: 1
subresources:
status: {}
You can checkout the CRDs on your cluster by running:
kubectl get crds
The sweet spot for CRDs is declarative state you want persisted in etcd and reconciled by a controller, like databases, certificates, canary rollouts, feature flags, tenancy objects — you name it. The CRD lives under your API group and versions (for example, widgets.platform.example.com) and exposes endpoints under /apis/<group>/<version>/<plural>. As soon as the CRD is accepted, kubectl get widgets works, discovery starts advertising your type, and kubectl explain widgets is backed by the OpenAPI you supplied. None of that requires custom server code; it’s all driven by the CRD spec and schema.
Once you’re structural, you can lean on OpenAPI-driven validation and defaulting. Defaulting is simple, add default to properties in your schema and the apiserver will materialize those defaults on create, the same way it does for built-in types. Validation starts with JSON Schema keywords but gets serious with CEL (Common Expression Language) rules: you attach x-kubernetes-validations expressions to fields and the apiserver evaluates them in-process. That means you can express constraints like “spec.size must be one of {S,M,L} unless spec.mode=debug” without maintaining a validating webhook. When you do need cross-field checks, CEL shines; it runs at admission time and blocks bad objects before they get stored in the Etcd. Here’s an example of a CRD with a set of CEL rules (some fields are dropped to reduce noise):
apiVersion: apiextensions.k8s.io/v1
kind: CustomResourceDefinition
spec:
versions:
- schema:
openAPIV3Schema:
properties:
spec:
properties:
size:
type: string
enum: ["S","M","L","XL"]
replicas:
type: integer
minimum: 0
opaqueConfig:
type: object
x-kubernetes-preserve-unknown-fields: true
# CEL rules
x-kubernetes-validations:
- rule: "!(self.size == 'XL' && !has(self.opaqueConfig))"
message: "XL requires an opaqueConfig block"
- rule: "!(self.size == 'S' && self.replicas > 1)"
message: "size=S cannot have more than 1 replica"
Versioning is first-class for CRDs (or really any API in Kubernetes). You declare one or more served versions and exactly one storage version. The API server persists objects in the storage version but will accept and return any served version, converting on the fly. If you later need to evolve the shape (in a backwards incompatible way) — rename fields, split enums, change nesting — you add a new served version and wire conversion. For built-ins, conversion functions live in compiled code; for CRDs, you attach a conversion webhook or rely on structural drop/rename behavior where applicable. We might discuss versioning further in a future blog post, but for now, checkout the upstream docs.
Most production CRDs also come with admission. CEL rules get you a long way, but you’ll often add validating and mutating webhooks for logic that’s hard to express declaratively — RBAC lookups, external constraints, or defaults that depend on context. Keep in mind CEL runs in-process and is operationally cheaper than webhooks; reach for webhooks when you truly need them. Meanwhile, RBAC for CRDs is business as usual: you grant verbs on your <plural>.<group> just like deployments.apps. Nothing special, but don’t forget watch permissions—controllers that rely on informers need list and watch to work efficiently.
On the controller side, CRDs slot neatly into the same client-go machinery as built-ins. Informers/cache watch your CRs and feed your reconciliation loop, you update status via the status endpoint, and you handle finalizers to implement clean deletion semantics. Just like any other resource, when a user deletes a Widget, the API server sets a deletion timestamp; your controller sees it, runs teardown against external systems, clears the finalizer, and only then does the object actually disappear. This keeps state in sync and prevents leaking cloud resources.
Finally, a word on UX and operability. CRDs publish OpenAPI to discovery, which means kubectl explain will surface your field docs and types; spend the extra minute annotating descriptions because users will see them. Use additionalPrinterColumns to make kubectl get useful at a glance instead of forcing everyone to pipe -o json | jq.
I won’t go through creating a CRD from scratch or using tools like kubebuilder, since that deserves it’s own separate post. For more information about CRDs, feel free to refer to the upstream docs.
Extension API Servers
This is where things get interesting — if you’re doing something wild, you may need to reach for an extension API server. With that you’ll get:
- You deploy a standalone API server (think: a Go app wired up with
k8s.io/apiserver) that implements its own REST endpoints. - Register it via an
APIServiceobject and kube-apiserver happily proxies requests to/apis/mygroup/v1/...onward. - You’re fully in control — storage, discovery, and behavior.
I personally think that CRDs are generally the right tool for the job — but there are moments when you absolutely want to break out of that declarative mold and do something more. That’s when I reach for an extension API server. The key realization is that Kubernetes doesn’t care whether an API it’s serving lives in the core binary or somewhere else entirely — it can proxy requests to other servers. In effect, the main kube-apiserver front-ends any number of extension servers via the aggregation layer. This tiny piece of architecture grants you massive flexibility.
API Server Chain
Let’s talk about how the API server chain is put together in upstream Kubernetes code, and what “delegation” and “aggregation” really mean in practice.
When you run kube-apiserver, what’s actually executing is a chain of (generic) APIServer implementations. Each layer is responsible for handling a subset of APIs and delegates to the next if it can’t serve the request. This is how different parts of the API are layered together:
- Core API server (
kube-apiserver): owns built-in APIs (/api/v1,/apis/apps/v1, etc.). - Extensions: admission chain, apiextensions-apiserver (which serves CRDs). CRDs ride on the apiextensions-apiserver, which is just another API server in the delegation chain.
- Aggregator: aggregation layer (which forwards requests for registered
APIServiceobjects).
Each piece has a “delegation target”, and when it doesn’t recognize a path, it just delegates down the chain until one does. Not-so-surprisingly, the last link in the chain, is a notFound handler that basically returns a 404.
Here’s how the code looks:
// CreateServerChain creates the apiservers connected via delegation.
func CreateServerChain(config CompletedConfig) (*aggregatorapiserver.APIAggregator, error) {
notFoundHandler := notfoundhandler.New(config.KubeAPIs.ControlPlane.Generic.Serializer, genericapifilters.NoMuxAndDiscoveryIncompleteKey)
apiExtensionsServer, err := config.ApiExtensions.New(genericapiserver.NewEmptyDelegateWithCustomHandler(notFoundHandler))
if err != nil {
return nil, err
}
crdAPIEnabled := config.ApiExtensions.GenericConfig.MergedResourceConfig.ResourceEnabled(apiextensionsv1.SchemeGroupVersion.WithResource("customresourcedefinitions"))
kubeAPIServer, err := config.KubeAPIs.New(apiExtensionsServer.GenericAPIServer)
if err != nil {
return nil, err
}
// aggregator comes last in the chain
aggregatorServer, err := controlplaneapiserver.CreateAggregatorServer(config.Aggregator, kubeAPIServer.ControlPlane.GenericAPIServer, apiExtensionsServer.Informers.Apiextensions().V1().CustomResourceDefinitions(), crdAPIEnabled, apiVersionPriorities)
if err != nil {
// we don't need special handling for innerStopCh because the aggregator server doesn't create any go routines
return nil, err
}
return aggregatorServer, nil
}
Aggregation
Aggregation is about exposing multiple API servers under one unified API surface. The API Aggregator runs inside the kube-apiserver process, watching APIService objects. Each APIService tells it:
- “For group X, version Y”
- “Forward requests to this Service in the cluster”
When a client requests /apis/metrics.k8s.io/v1/..., the aggregator sees an APIService for metrics.k8s.io/v1, proxies the request to the metrics-server, and returns the response. To the client, it looks like a native Kubernetes API group. Extension API servers are external processes, but kube-apiserver fronts them through the aggregator.
The aggregator also participates in discovery: when you hit /apis, it merges built-in groups and all aggregated ones into a single response. That’s why CRDs, built-ins, and aggregated APIs all appear side by side.
Under the hood, Kubernetes uses APIService objects in the apiregistration.k8s.io/v1 API group to broker the integration. You declare which service in your cluster handles the path, drop a CA bundle in there so TLS works end-to-end, and Kubernetes will handle proxying, discovery, and even RBAC headers for you. Here’s how the APIService for metrics server looks like:
apiVersion: apiregistration.k8s.io/v1
kind: APIService
metadata:
name: v1beta1.metrics.k8s.io
spec:
group: metrics.k8s.io
insecureSkipTLSVerify: true
service:
name: metrics-server
namespace: kube-system
port: 443
version: v1beta1
You can checkout the APIService objects in your cluster by running:
kubectl get apiservice
APIServiceobjects are interesting. There is one even for every built-in Kubernetes version/group (e.g.v1.apps) or version/groups added by your CRDs. You can mess with your cluster by editing the APIServices responsible for built-in version/groups. In a Kubernetes cluster, you’ll notice that there are many of them without aservicefield. So, they’re not proxying requests to an extension API server? No. Then why are they there at all? Well, good question. Short answer: Mostly for the sake of discovery. We might exploreAPIServiceobjects in a later blog post.
The beauty of this pattern is how it unlocks real-time, imperative behavior in your API. Think of things like /exec, /logs, or /portforward. Those subresources exist in the core server for Pods, but you can add your own for your types. Maybe you want /restart, /console, or /vnc on a virtual machine resource. Kubernetes doesn’t just let you define those endpoints, it actually handles authn, RBAC, and discovery for them the same way it does for built-in ones.
Do you want your API server to hold state in Postgres instead of etcd? No problem. CRDs lock you into the main API server, but an extension API server can choose any storage or backend.
That said, construction of an extension API server isn’t trivial work. You’re building a Kubernetes-grade API yourself — including storage, schema generation, OpenAPI documentation, authn/authz delegation, and REST handling. You’ve got to wire up delegated authentication so that users aren’t re-authenticated inside your server, and propagate their user context with headers. You’ve got to handle subject access review to authorize actions. All this plumbing the core server does for built-ins — now you own it or glue it in with libraries like k8s.io/apiserver.
There’s operational complexity too. Your extension API server must respond swiftly — discovery requests have tight latency budgets. A slow extension server can bog down the entire control plane. And if your server crashes, suddenly namespace deletion hangs because Kubernetes can’t clean up resources it still thinks exist. That’s why this approach is powerful, but only when you need that power and can own the operational burden.
Finally, if you still crave patterns like client generation or simplified scaffolding, there’s a starting point for setting up an extension API server, the apiserver-builder project (still alpha, fyi) and the classic sample‑apiserver example. These give you a starting point, code generation, internal type scaffolding, and wiring, but from there you own the rest.
At the end of the day, a custom aggregated API server is Kubernetes saying: “You really know what you’re doing, here’s how you plug into my front door.” Use it when CRDs aren’t enough, and you need full control — not lightly, but when it matters, it’s glorious.
How CRDs Differ from Extension API Servers
I tend to think of CRDs as the “easy path” — the one you take when you want to get up and running fast, but still want resources that feel native. You define your schema, slap it into etcd via the kube-apiserver, and suddenly the magic happens. Discovery works, kubectl understands it, you can do validation, defaulting, server-side apply, status subresources, etc. All without spinning up another service or writing boilerplate server code. The kube-apiserver handles the heavy lifting and your controller just watches and reconciles. That’s powerful.
Contrast that with an extension API server. This is the more deliberate, hands-on route. Instead of just defining a resource type, you build a whole API server yourself — schema, storage, discovery, authorization, OpenAPI spec, everything. You host it in the cluster, register an APIService so the main API server delegates requests down to you, and to clients it looks exactly like any other Kubernetes API group. Probably the best example is metrics-server: it appears at /apis/metrics.k8s.io/, but it’s served by a separate process, not the core kube-apiserver. That separation gives you full freedom — to choose storage layers, to do imperative logic, to pull data from external systems, or to serve ephemeral, RPC-style endpoints. But it also means you’ve inherited all the responsibility.
Ease of use and operational gravity pull CRDs into most scenarios. They’re inherently less brittle — no extra service to monitor, no SLA to uphold, no TLS wiring for API discovery, TLS cert rotation, or proxy latency to optimize. The kube-apiserver keeps everything in a single control plane vessel.
Validation and defaulting feel slightly different too. With CRDs, you declare defaults using OpenAPI schema keywords, and validation can be as powerful as you like with CEL and schema. You get defaults auto-magically, and validation is managed inside the API server. Extension servers need to own that logic — either co-opt defaulting via openAPI generated code or run your own webhook or inline validators. The flip side is limitless flexibility, but more lines of code.
At the end of the day, most teams reach for CRDs first. They’re faster to try out, easier to maintain, and integrate cleanly with the Kubernetes API model. But extension servers are invaluable when your use case stretches the API pattern — when you need custom subresources, external data, RPC behavior, non-etcd persistence, or any logic the API server just can’t or shouldn’t handle. It’s more work, but if you need that control, a custom API server gives you full ownership.
Closing Thoughts
Stepping back, what actually excites me about both CRDs and extension API servers is not merely how they let you extend Kubernetes — it’s that they let you redefine what Kubernetes itself is for your team. You get to choose whether your cluster is a place of purely declarative abstractions — objects stored in etcd and reconciled by controllers — or whether it’s a programmable hub that can talk to external systems, compute responses, and elevate your own platform’s capabilities.
CRDs are fast to author, easy to test, and the Kubernetes ecosystem supports them in all the ways that matter — kubectl, server-side apply, schema validation, and the status subresource from day one. Controllers built with controller-runtime slot right in, and SSA ensures your resources merge cleanly. For 90 % of use cases — operators, declarative multi-step workflows, custom workloads — a well-designed CRD is the cleanest, simplest path.
But there will come a time when CRDs feel like a frame that’s too small for the picture you have in your head. Maybe you want Rails-style introspection, a /restart action on your custom VM resource, or real-time, RPC-esque responses when you query an external system. That’s when you lean in—to build an extension API server. It’s more work, sure, and there’s TLS wiring, discovery latency, authentication proxying, and all the usual ops concerns. But when you get it right, your API is indistinguishable from a built-in one—and it can do things no CRD ever could. It becomes a true extension of the control plane.
Thanks for taking the time to read this. If you’d like to continue the conversation or share your own experiences, feel free to connect with me on LinkedIn, or checkout my personal website.
Further Reading
Container Network Interface (CNI) in Kubernetes: An Introduction — We’re gonna learn about the Container Network Interface (CNI) and CNI plugins, what they’re supposed to do, and how…
Seamless Cluster Creation & Management: Canonical Kubernetes v1.32 Stable Is Released — With the release of the shiny new 1.32 stable version, the Canonical Kubernetes solidifies itself as a dependable…
How Canonical Kubernetes CAPI Providers Handle In-Place Upgrades — In-Place Upgrades Design and Implementation in Canonical Kubernetes Cluster API Providers
gRPC Name Resolution & Load Balancing on K8s: Everything you need to know (and probably a bit more) — Load balancing gRPC requests on Kubernetes can be challenging. In this blog we tried to deep dive into the…