This is the multi-page printable view of this section.
Click here to print.
Return to the regular view of this page.
Concepts
The Goten concepts a service author needs to know.
Background on how Goten works underneath the guide. You do not need all of this
to build a service, but it explains the machinery you are building on and is
worth reading when something does not behave as you expect.
In this section
1 - Meta service as service registry
Understanding the role of Meta service as a service registry.
To build a multi-service framework, we first need a special service,
that provides service registry offers. Using it, we must be able to
discover:
- List of existing Regions
- List of existing Services
- List of existing Resources per Service
- List of existing regional Deployments per Service.
This is provided by the meta.goten.com Service, in the Goten repository,
directory meta-service. It follows the typical structure of any service,
but has no cmd directory or fixtures, as Goten provides only basic parts.
The final implementation is in the edgelq repository, see directory meta.
SPEKTRA Edge version of meta contains an old version of the service, v1alpha2,
which is obsolete and irrelevant to this document. For this purpose, ignore
v1alpha2 elements.
Still, the resource model for Meta service resides in the Goten repository,
see normal protobuf files. For Goten, we made the following design decisions,
this reflects fields we have in protobuf files (you can and should see).
- List of regions in meta service must show a list of all possible
regions where services can be deployed, not necessarily where are
deployed.
- Each Service must be fairly independent. It must be able to specify
its global network endpoint where it is reachable. It must display
a list of API versions it has. For each API version, it must tell
which services it imports, and which versions of them. It must tell
what services it would like to use as a client too (but not import).
- Every Deployment describes an instance of a service in a region.
It must be able to specify its regional network endpoint and tell
which service version it operates on (current maximum version). It is
assumed it can support lower versions too. Deployments for a single
service do not need to upgrade at once to the new version, but it’s
recommended to not wait too long.
- Deployments can be added to a Service dynamically, meaning, service
owners can expand by just adding new Deployment in Meta service.
- Each Service manages its multi-region setup. Meaning: Each Service
decides which region is “primary” for them. Then list of Deployment
resources describes what regions are available.
- Each region manages its network endpoints, but it is recommended to
have the same domain for global and regional endpoints, and each
regional endpoint has a region ID as part of a subdomain, before
the main part.
- For Service A to import Service B, we require that Service B is
available in all regions where Service A is deployed. This should
be the only limitation Services must follow for multi-region setup.
All those design decisions are reflected in protobuf files, and server
implementation (custom middlewares), see in goten repository,
meta-service/server/v1/ custom middlewares, they are fairly simple.
For SPEKTRA Edge, design decisions are that:
- All core SPEKTRA Edge services (iam, meta adaptation, audit, monitoring,
etc.) are always deployed to all regions and are deployed together.
- It means, that 3rd party services can always import any SPEKTRA Edge core
service because it is guaranteed to be in all regions needed by 3rd
party.
- All core SPEKTRA Edge services will point to the same primary region.
- All core SPEKTRA Edge services will have the same network domain:
iam.apis.edgelq.com, monitoring.apis.edgelq.com, etc.
If you replace the first word with another, it will be valid.
- If core SPEKTRA Edge services are upgraded in some regions, then they
will be upgraded at once.
- All core SPEKTRA Edge services will be public: Anyone authenticated will
be able to read its roles, permissions, and plans, or be able to import
them.
- All 3rd party services will be assumed to be users of core SPEKTRA Edge
services (no cost if no actual use).
- Service resources can be created by a ServiceAccount only. It is
assumed that it will be managing this Service.
- Service will belong to a Project, where ServiceAccount who created
it belongs.
Users may think of core edgelq services as a service bundle. Most of
these SPEKTRA Edge rules are declarations, though the deployment workflows are
expected to enforce them regardless. The decision, that all 3rd parties are
considered users of all core SPEKTRA Edge services, and that each Service must
belong to some project, is reflected in additional custom middleware we
have for meta service in the edgelq repository, see file
meta/server/v1/service/service_service.go. In this extra middleware,
executed before custom middleware in the goten repository
(meta-service/server/v1/service/service_service.go), we are
adding core SPEKTRA Edge to the used services array. We also assign
a project-owning Service. This is where the management of ServiceAccounts
is, or where usage metrics will go.
This concludes Meta service workings, where we can find information about
services and relationships between them.
2 - EnvRegistry as service discovery
Understanding the role of EnvRegistry module in Meta service.
Meta service provides API allowing inspection global environment, but
we also need a side library, called EnvRegistry:
- It must allow a Deployment to register itself in a Meta service,
so others can see it.
- It must allow the discovery of other services with their deployments
and resources.
- It must provide a way to obtain real-time updates of what is happening
in the environment.
Those three items above are the responsibilities of EnvRegistry module.
In the goten repo, this module is defined in the
runtime/env_registry/env_registry.go file.
As of now, it can only be used by server, controller, and
db-controller runtimes. It may be beneficial for client runtimes
someday probably, but we will opt out from “registration” responsibility
because the client is not the part of the backend, it cannot self-register
in Meta service.
One of the design decisions regarding EnvRegistry is that it must block
till initialization is completed, meaning:
- User of EnvRegistry instance must complete self-registration
in Meta service.
- EnvRegistry must obtain the current state of services and
deployments.
Note that no backend service works in isolation, as part of the Goten design,
it is essential that:
- any backend runtime knows its surroundings before executing
its tasks.
- all backend runtimes must be able to see other services and
deployments, which are relevant for them.
- all backend runtimes must initialize and run the EnvRegistry
component and it must be one of the first things to do in
the
main.go file.
This means, that the backend service, if it cannot successfully pass
initialization, will be blocked from any useful work. If you check
all run functions in EnvRegistry, you should see they lead to
the runInBackground function. It runs several goroutines, but
then it waits for a signal showing all is fine. After this,
EnvRegistry can be safely used to find other services, and
deployments, and make networking connections.
This also guarantees that Meta service contains relevant records
for services, in other words, EnvRegistry registration initializes
regions, services, deployments, and resources. Note,
however:
- The region resources can be created/updated by meta.goten.com
service only. Since meta is the first service, it is responsible
for this resource to be initialized.
- The service resource is created by the first deployment of
a given service. So, if we release custom.edgelq.com for the
first time, in the first region, it will send a CreateService
request. The next deployment of the same service, in the next
region, will just send UpdateService. This update must have
a new MultiRegionPolicy, where field-enabled regions contain
a new region ID.
- Each deployment is responsible for its deployment resource
in Meta.
- All deployments for a given service are responsible for
Resource instances. If a new service is deployed with
the server, controller, and db-controller pods, then they may
initially be sending clashing create requests. We are fine with
those minor races there, since transactions in Meta service, coupled
with CAS requests made by EnvRegistry, ensure eventual consistency.
Visit the runInit function, which is one of the goroutines of
EnvRegistry executed by runInBackground. It contains procedures
for registration of Meta resources finishes after a successful run.
From this process, another emerging design property of EnvRegistry
is that it is aware of its context, it knows what Service and Deployment
it is associated with. Therefore, it has getters for self Deployment and
Service.
Let’s stay for a while in this run process, as it shows other goroutines
that are run forever:
- One goroutine keeps running runDeploymentsWatch
- Second goroutine keeps running runServicesWatch
- The final goroutine is the main one, runMainSync
We don’t need real-time watch updates of regions and resources, we need
services and their regional deployments only. Normally watch requires
a separate goroutine, and it is the same case here. To synchronize actual
event processing across multiple real-time updates, we need a “main
synchronization loop”, which unites all Go channels.
In the main sync goroutine, we:
- Process changes detected by runServicesWatch.
- Process changes detected by runDeploymentsWatch.
- Catch initialization signal from the runInit function, which
guarantees information about our service is stored in Meta.
- Attachment of new real-time subscribers. When they attach, they
must get a snapshot of past events.
- Detachment of real-time subscribers.
As of additional note: since EnvRegistry is self-aware, it gets
only Services and Deployments that are relevant. Those are:
- Services and Deployments of its Service (obviously)
- Services and Deployments that are used/imported by the current Service
- Services and Deployments that are using the current Service
The last two parts are important, it means that EnvRegistry for top
service (like meta.goten.com) is aware of all Services and Deployments.
Higher levels will see all those below or above them, but they won’t be able
to see “neighbors”. The higher the tree, there will be fewer services above,
and more below, but the proportion of neighbors will be higher and higher.
It should not be a problem, though, unless we reach the scale of thousands
of Services, core SPEKTRA Edge services will however be more pressured than all
upstream ones for various reasons.
In the context of SPEKTRA Edge, we made additional implementation decisions,
when it comes to SPEKTRA Edge platform deployments:
-
Each service, except meta.goten.com itself, must connect to the
regional meta service in its EnvRegistry.
For example, iam.edgelq.com in us-west2, must connect to Meta
service in us-west2. Service custom.edgelq.com in eastus2 must
connect to Meta service in eastus2.
-
Server instance of meta.goten.com must use local-mode
EnvRegistry. The reason is, that it can’t connect to itself
via API, especially since it must succeed in EnvRegistry
initialization before running its API server.
-
DbController instance of meta.goten.com is special, and shows
the asymmetric nature of SPEKTRA Edge core services regarding regions.
As a whole, core SPEKTRA Edge services point to the same primary region,
any other is secondary. Therefore, DbController instance of
meta.goten.com must:
- In the primary region, connect to the API server of
meta.goten.com in the primary region (intra-region)
- In the secondary region, connect to the API server of
meta.goten.com in the primary region (the secondary
region connects to the primary).
Therefore, when we add a new region, the meta-db-controller in
the secondary region registers itself in the primary region
meta-service. This way primary region gets the awareness of
the next region’s creation. The choice of meta-db-controller for
this responsibility has more for it, Meta-db-controller will be
responsible for syncing the secondary region meta database from
the primary one. This will be discussed in the following section
of this guide. For now, we just mentioned conventions where
EnvRegistry must source information from.
3 - Resource metadata
Understanding the resource metadata for the service synchronization
As a protocol, Goten needs to have protocol-like properties. One of
the thems is the requirement that resource types of all Services managed
by Goten must contain metadata objects. It was already mentioned multiple
times, but let’s put a link to the Meta object again
https://github.com/cloudwan/goten/blob/main/types/meta.proto.
Resource type managed by Goten must satisfy interface methods
(you can see in the Resource interface defined in the
runtime/resource/resource.go file):
GetMetadata() *meta.Meta
EnsureMetadata() *meta.Meta
There is, of course, the option to opt-out, interface Descriptor has
method SupportsMetadata() bool. If it returns false, it means
the resource type is not managed by Goten, and will be omitted from
the Goten design! However, it is important to recognize if resource
type is subject to this design or not, and how we can do this, including
programmatically.
To summarize, as protocol, Goten requires resources to satisfy this
interface. It is important to note what information is stored in resource
metadata in the context of the Goten design:
-
Field syncing of type SyncingMeta must always describe which region
owns a resource, and which regions have read a copy of it. SyncingMeta
must be always populated for each resource, regardless of type.
-
Field services of type ServicesInfo must tell us which service
owns a given resource, and a list of services for which this resource
is relevant. Unlike syncing, services may not be necessarily populated,
meaning that Service-defining resource type is responsible for explaining
how it works in this case. In the future probably it may slightly
change:
If services is not populated at the moment of resource save, it will
point to the current service as owning, and allowed services will be
a one-element array containing the current service too. This in fact
should be assumed by default, but it is not enforced globally, which
we will explain now.
First, service meta.goten.com always ensures that the services
field is populated for the following cases:
- Instances of meta.goten.com/Service must have ServicesInfo where:
- Field
owning_service is equal to the current service itself.
- Field
allowed_services contains the current service, all
imported/used services, AND all services using importing
this service! Note that this may be dynamically changing, if
a new service is deployed, it will update the ServicesInfo fields
of all services it uses/imports.
- Instances of meta.goten.com/Deployment and
meta.goten.com/Resource must have their ServicesInfo
synchronized with parent meta.goten.com/Service instance.
- Instances of meta.goten.com/Region do not have ServicesInfo
typically populated. However, in the SPEKTRA Edge context, we have
a public RoleBinding that allows all users to read from this
collection (but never write). Because of this private/public
nature, there was no need to populate service information there.
Note that this implies that service meta.goten.com is responsible for
syncing ServicesInfo of meta.goten.com/Deployment and
meta.goten.com/Resource instances. It is done by a controller
implemented in the Goten repository: meta-service/controller
directory. It is relatively simple.
However, while meta.goten.com can detect what ServicesInfo should be
populated, this is often not the case at all. For example, when service
iam.edgelq.com receives a request CreateServiceAccount, it does not
know necessarily for whom this ServiceAccount is at all. Multiple services
may be owning ServiceAccount resources, therefore, but the resource type
itself does not have a dedicated “service” field in its schema. The only
way services can annotate ServiceAccount resources is by providing necessary
metadata information. Furthermore, if some custom service wants to make
the ServiceAccount instance available for others services to see, it may
need to provide multiple items to the allowed_services array. This should
explain that service information must be determined at the business logic
level. For this reason, it is allowed to have empty service information,
but in many cases, SPEKTRA Edge will enforce their presence, where business
logic requires it.
Then, the situation for the other meta field, syncing, is much easier.
Value can be determined on the schema level. There already is instruction
in the multi-region design section of the developer guide.
Regions setup always can be defined based on resource name only:
- If it is a regional resource (has a
region/ segment in the name),
it strictly tells which region owns it. The list of regions that
get a read-only copy is decided on below resource name properties
below.
- If it contains a well-known policy-holder in the name, then
the policy-holder defines what regions get a read copy. If
the resource is non-regional, then MultiRegionPolicy also tells
what region owns it (default control region).
- If the resource is not subject to MultiRegionPolicy (like Region,
or User in iam.edgelq.com), then it is a subject of
MultiRegionPolicy defined in the relevant meta.goten.com/Service
instance (for this service).
Now the trick is: All policy-holder resources are well-known. Although we
try not to hardcode anything anywhere, Goten provides utility functions
for detecting if a resource contains a MultiRegionPolicy field in its
schema. This also must be defined in the Goten specification. By detecting
what resource types are policy-holders, Goten can provide components that
can easily extract regional information from a given resource by its name
only.
Versioning information does not need to be specified in the resource body.
Having instance, it is easily possible to get Descriptor instance, and
check API version. All schema references are clear in this regard too, if
resource A has a reference field to resource B, then from the reference
object we can get the Descriptor instance of B, and get the version.
The only place where it is not possible, are meta owner references.
Therefore, in the field metadata.owner_references, an instance of
each must contain the name, owning service, API version, and region
(just in case it is not provided in the name field). When talking about
the meta references, it is important to mention other differences
compared to schema-level references:
- schema references are owned by a Service that owns resources
with references.
- meta owner references are owned by a Service to which references
are pointing!
This ownership has implication: when Deployment D1 in Service S1 upgrades
from v1 to v2 (for example), and there is some resource X in Deployment
D2 from Service S2, and this X has the meta owner reference to some
resource owned by D1, then D1 will be responsible for sending an Update
request to D2, so meta owner reference is updated.
4 - Multi-region policy store
Understanding the design of the multi-region policy store.
We mentioned MultiRegion policy-holder resources, and their importance
when it comes to evaluating region syncing information based on resource
name. There is a need to have a MultiRegion PolicyStore object, that
for any given resource name returns a managing MultiRegionPolicy object.
This object is defined in the Goten repository, file
runtime/multi_region/policy_store.go. This file is important for this
design and worth remembering. As of now, it returns a nil object for global
resources though, the caller should in this case take MultiRegionPolicy
from the EnvRegistry component from the relevant Service.
It uses a cache that accumulates policy objects, so we should normally
not use any IO operations, only initially. We have watch-based invalidation,
which allows us to have a long-lived cache.
We have some code-generation that provides us functions needed to
initialize PolicyStore for a given Service in a given version, but
the caller is responsible for remembering to include them (All those
main.go files for server runtimes!).
In this file, you can also see a function that sets/gets MultiRegionPolicy
from a context object. In multi-region design, it is required from a server
code, to store the MultiRegionPolicy object in a context if there will be
updates to the database!
5 - Goten organization
Understanding the Goten directory structure and libraries.
In the SPEKTRA Edge repository, we have directories for each service:
edgelq/
applications/
audit/
devices/
iam/
limits/
logging/
meta/
monitoring/
proxies/
secrets/
ztp/
All names of these services end with the .edgelq.com suffix, except meta.
The full name of the meta service is meta.goten.com. The reason is that
the core of this service is not in the SPEKTRA Edge repository, it is in the
Goten repository:
goten/
meta-service/
This is where meta’s api-skeleton is, Protocol buffers files, almost all
the code-generated modules, and server implementation. The reason why
we talk about meta service first is quite important, because it also teaches
the difference between SPEKTRA Edge and Goten.
Goten is called a framework for SPEKTRA Edge, but this framework has two main
tool sets:
-
Compiler
It takes the schema of your service, and generates all the boilerplate
code.
-
Runtime
Runtime libraries, which are referenced by generated code, are used
heavily throughout all services based on Goten.
Goten provides its schema language on top of Protocol Buffers: It introduces
the concept of Service packages (with versions), API groups, actions,
resources. Unlike raw Protocol Buffers, we have a full-blown schema with
references that can point across regions and services, and those services
can also vary in versions. Resources can reference each other, services can
import each other.
Goten balances between code-generation and runtime libraries operating on
“resources” or “methods”. It is usually the tradeoff between performance,
type safety, code size, maintainability, and the readability.
If you look at the meta-service, you will see that it has four resource
types:
- Region
- Service
- Resource
- Deployment
This is pretty much exactly what Goten provides. To be a robust framework,
to provide on its promises, Multi-region, multi-service, and multi-version,
Goten needs a concept of a service that contains information about regions,
services, etc.
SPEKTRA Edge provides various services on a higher level, but Goten provides
the baseline for them and allows relationships between them. The only
reason why the “meta” directory exists also in SPEKTRA Edge, is because
the meta service also needs extra SPEKTRA Edge integration like authorization
layer. In the SPEKTRA Edge repo, we have additional components added to meta,
and finally, we have meta main.go files. If you look at the files meta
service has in the SPEKTRA Edge repo (for the v1 version, not v1alpha2), you
will understand that the edgelq version wraps up what goten provides. There is
also a full “v1alpha2” service, which predates the move of meta into Goten. At
that time SPEKTRA Edge overrode much of the functionality Goten provides.
Goten Directory Structure
As a framework, goten provides:
- Modules related to service schema and prototyping (API skeleton and
proto files).
- Compilers that generate code based on schema
- Runtime libraries linked during the compilation
For schema & prototyping, we have directories:
-
schemas
This directory contains generated JSON schema for api-skeleton files.
It is generated based on file annotations/bootstrap.proto.
-
annotations
Protobuf is already a kind of language for building APIs, but Goten is
said to provide a higher level one. This directory contains various
proto options (extra decorations), that enhance standard protobuf
language. There is one exceptional file though: bootstrap.proto,
which DOES NOT define any options, instead it describes
the api-skeleton schema in protobuf. The file in the schemas directory
is just a compilation of this file. The annotations directory contains
generated Golang code describing those proto options, which you do not
normally need to read.
-
types
Contains set of reusable protobuf messages that are used in services
using Goten, for example, “Meta” object (file types/meta.proto) is
used in almost every resource type. The difference between annotations
and types is that, while annotations describe options we can attach
to files/proto messages/fields/enums etc., types contain just reusable
objects/enums. Apart from that, each proto file contains compiled
Golang objects in the relevant directory.
-
contrib/protobuf/google
This directory avoids depending on the full protobuf distribution from
Google by vendoring only the parts SPEKTRA Edge uses. Subdirectory api
maps to annotations, and type to types. One file does not follow that
mapping: distribution.proto belongs more with the type directories than
with api. It does not appear to be used, and may be removable.
Contributors do still download some protocol buffers separately (see the
scripts directory). That downloaded library is more lightweight and does
not contain the types vendored here.
All the above directories can be considered a form of Goten-protobuf language that you should know from the developer guide.
For compilers (code-generators), we have directories:
-
compiler
Each subdirectory (well, almost) contains a specific compiler that
generates some set of files that Goten as a whole generates. For
example compiler/server generates server middleware.
-
cmd
Goten does not come with any runtime on its own. This directory
provides main.go files for all compilers (code-generators) Goten has.
Compilers generate code you should already know from the developer guide
as well.
Runtime libraries have just a single directory:
-
runtime
Contains various modules for clients, servers, controllers… Each will
be talked about separately in various topic-oriented documents.
-
Compiled types
types/meta/, and types/multi_region_policy may be considered part of
the runtime, they map to types objects. You may say, that while resource
proto schema imports goten/types/meta.proto, generated code will refer
to Go package goten/types/meta/.
In the developer guide, we had brief mentions of some base runtime types,
but we were treating them as black boxes, while in this document set, we
will dive in.
Other directories in Goten:
-
example
Contains some typical services developed on Goten, but without SPEKTRA Edge.
The current purpose of them is only to run some integration tests though.
-
prototests
Contains just some basic tests over base extended types by Goten, but
does not delve as deep as tests in the example directory.
-
meta-service
It contains full service of meta without SPEKTRA Edge components and main
files. It is supposed to be wrapped by Goten users, SPEKTRA Edge in our case.
-
scripts
Contains one-of scripts for installing development tools, reusable
scripts for other scripts, or regeneration script that regenerates
files from the current goten directory (regenerate.sh).
-
src
This directory name is the most confusing here. It does not contain
anything for the framework. It contains generated Java code of
annotations and types directories in Goten. It is generated for
the local pom.xml file. This Java module is just an import dependency
for Goten, so Java code can use protobuf types defined by Goten. We have
some Java code in the SPEKTRA Edge repository, so for this purpose, in
Goten, we have a small Java package.
-
tools
Just some dummy imports to ensure they are present in go.mod/go.sum
files in goten.
-
webui
A generic UI for Goten services. This is no longer maintained, as the
front-end teams now build specialized UIs rather than a generic one.
The remaining files worth noting:
-
pom.xml
This is for building a Java package containing Goten protobuf types.
-
sdk-config.yaml
This is used to generate the public goten-sdk repository, since goten
itself is private. It automates copying the public files from goten to
goten-sdk.
-
tools.go
Ensures the required dependencies are present in go.mod.
SPEKTRA Edge Directory Structure
SPEKTRA Edge is a home repository for all core SPEKTRA Edge services, and
an adaptation of meta.goten.com, meaning that its sub directories should
be familiar, and you should navigate their code well enough since they are
“typical” Goten-built services.
We have a common directory though, with some example elements
(more important):
-
api and rpc
Those directories contain extra protobuf reusable types. You will
most likely interact with api.ServiceAccount (not to confuse with
the iam.edgelq.com/ServiceAccount resource)!
-
cli_configv1, cli_configv2
The second directory is used by the cuttle CLI utility, and will be
needed for all cuttles for 3rd parties.
-
clientenv
Those contains obsolete config for client env, but its grpc dialers
and authclients (for user authentication) are still in use. Needs some
cleanup.
-
consts
It has a set of various common constants in SPEKTRA Edge.
-
doc
It wraps protoc-gen-goten-doc with additional functionality, to
display needed permissions for actions.
-
fixtrues_controller
It is the full fixtures controller module.
-
serverenv
It contains a common set for backend runtimes provided by SPEKTRA Edge
(typically server, but some elements are used by controllers too).
-
widecolumn
It contains a storage alternative to the Goten store, for some
advanced cases, we will have a different document design for this.
Other directories:
-
healthcheck
It contains a simple image that polls health checks of core SPEKTRA Edge
services.
-
mixins
It contains a set of mixins, they will be discussed via separate topics.
-
protoc-gen-npm-apis
It is a Typescript compiler for the frontend team, maintained by
the backend. You should read more about
compilers here
-
npm
It is where code generated by protoc-gen-npm-apis goes.
-
scripts
Set of common scripts, developers must learn to use primarily
regenerate-all-sh whenever they change any api-skeleton or
proto file.
-
src
Contains Java-generated code for the Monitoring Pipeline, which is being
phased out and will be documented separately.
In this section
5.1 - Goten server library
Understanding the Goten server library.
The server should more or less be already known from the developer guide.
We will provide some missing bits here only.
When we talk about servers, we can distinguish:
- gRPC Server instance that is listening on a TCP port.
- Server handler sets that implement some Service GRPC interface.
The following code snippet from IAM shows this:
grpcServer := grpcserver.NewGrpcServer(
authenticator.AuthFunc(),
commonCfg.GetGrpcServer(),
log,
)
v1LimMixinServer := v1limmixinserver.NewLimitsMixinServer(
commonCfg,
limMixinStore,
authInfoProvider,
envRegistry,
policyStore,
)
v1alpha2LimMixinServer := v1alpha2limmixinserver.NewTransformedLimitsMixinServer(
v1LimMixinServer,
)
schemaServer := v1schemaserver.NewSchemaMixinServer(
commonCfg,
schemaStore,
v1Store,
policyStore,
authInfoProvider,
v1client.GetIAMDescriptor(),
)
v1alpha2MetaMixinServer := metamixinserver.NewMetaMixinTransformerServer(
schemaServer,
envRegistry,
)
v1Server := v1server.NewIAMServer(
ctx,
cfg,
v1Store,
authenticator,
authInfoProvider,
envRegistry,
policyStore,
)
v1alpha2Server := v1alpha2server.NewTransformedIAMServer(
cfg,
v1Server,
v1Store,
authInfoProvider,
)
v1alpha2server.RegisterServer(
grpcServer.GetHandle(),
v1alpha2Server,
)
v1server.RegisterServer(grpcServer.GetHandle(), v1Server)
metamixinserver.RegisterServer(
grpcServer.GetHandle(),
v1alpha2MetaMixinServer,
)
v1alpha2limmixinserver.RegisterServer(
grpcServer.GetHandle(),
v1alpha2LimMixinServer,
)
v1limmixinserver.RegisterServer(
grpcServer.GetHandle(),
v1LimMixinServer,
)
v1schemaserver.RegisterServer(
grpcServer.GetHandle(),
schemaServer,
)
v1alpha2diagserver.RegisterServer(
grpcServer.GetHandle(),
v1alpha2diagserver.NewDiagnosticsMixinServer(),
)
v1diagserver.RegisterServer(
grpcServer.GetHandle(),
v1diagserver.NewDiagnosticsMixinServer(),
)
There, an instance called grpcServer is an actual GRPC Server instance
listening on a TCP port. If you dive into this implementation, you should
notice we are constructing an EdgelqGrpcServer structure. It may consist
of actually two port listening instances:
googleGrpcServer *grpc.Server, which is initialized with a set of
unary and stream interceptors, optional TLS.
websocketHTTPServer *http.Server, which is initialized only if
the websocket port was set. It delegates handling to
improbableGrpcwebServer, which uses googleGrpcServer.
This Google server is the primary one and handles regular gRPC calls.
The reason for the additional HTTP server is that we need to support
web browsers, which cannot support native gRPC protocol. Instead:
- grpcweb is needed to handle unary and server-streaming calls.
- websockets are needed for bidirectional streaming calls.
Additionally, we have REST API support…
We have this envoy proxy sidecar, a separate container running next to
the server instance. It handles all REST API, converting to native gRPC.
It converts grpcweb into native grpc too, but has issues with websockets.
For this reason, we added a Golang HTTP server with an improbable gRPC web
instance. This improbable grpc web instance can handle both grpcweb and
websockets, but we use it for websockets only, since it is missing from
envoy proxy.
In theory, an improbable web server would be able to handle ALL protocols,
but there is a drawback: For native gRPC calls will be less performant than
the native grpc server (and ServeHTTP is less maintained). It is recommended
to keep them separate, so we will stick with 2 ports. We may have some
opportunity to remove the envoy proxy though.
Returning to the googleGrpcServer instance, we have all stream/unary
interceptors that are common for all calls, but this does not implement
the actual interface we expect from gRPC servers. Each service version
provides a complete interface to implement. For example, see the IAMServer
interface in this file:
https://github.com/cloudwan/edgelq/blob/main/iam/server/v1/iam/iam.pb.grpc.go.
Those server interfaces are in files ending with pb.grpc.go.
To have a full server, we need to combine the GRPC Server instance for
SPEKTRA Edge (EdgelqGrpcServer), with, let’s make up some name for it:
A business logic server instance (set of handlers). In this iam.pb.grpc.go
file this business logic instance is iamServer. Going back to the main.go
snippet that is provided way above, we are registering eight business
logic servers (handler sets) on the provided *grpc.Server instance.
As long as paths are unique across all, it is fine to register as many as
we can. Typically, we must include primary service for all versions, then
all mixins in all versions.
Those business logic servers provide code-generated middleware, typically
executed in this order:
- Multi-region routing middleware (may redirect processing somewhere else,
or split across many regions).
- Authorization middleware (may use a local cache, or send a request to
IAM to obtain fresh role bindings).
- Transaction middleware (configures access to the database, for snapshot
transactions and establishes new session).
- Outer middleware, which provides validation, and common outer operations
for certain CRUD requests. For example, for update calls, it will ensure
the resource exists and apply an update mask to achieve the final resource
to save.
- Optional custom middleware and server code - which are responsible for
final execution.
Transaction middleware also may repeat execution of all internal middleware
- core server, if the transaction needs to be repeated.
There are also “initial handlers” in generated pb.grpc.go files. For
example, see this file:
https://github.com/cloudwan/edgelq/blob/main/iam/server/v1/group/group_service.pb.grpc.go.
For example, you can see _GroupService_GetGroup_Handler as example for
unary, and _GroupService_WatchGroup_Handler as an example for streaming
calls.
It is worth mentioning how interceptors play with middleware and these
“initial handlers”. Let’s copy and paste interceptors from the current
edgelq/common/serverenv/grpc/server.go file:
grpc.StreamInterceptor(grpc_middleware.ChainStreamServer(
grpc_ctxtags.StreamServerInterceptor(),
grpc_logrus.StreamServerInterceptor(
log,
grpc_logrus.WithLevels(codeToLevel),
),
grpc_recovery.StreamServerInterceptor(
grpc_recovery.WithRecoveryHandlerContext(recoveryHandler),
),
RespHeadersStreamServerInterceptor(),
grpc_auth.StreamServerInterceptor(authFunc),
PayloadStreamServerInterceptor(log, PayloadLoggingDecider),
grpc_validator.StreamServerInterceptor(),
)),
grpc.UnaryInterceptor(grpc_middleware.ChainUnaryServer(
grpc_ctxtags.UnaryServerInterceptor(),
grpc_logrus.UnaryServerInterceptor(
log,
grpc_logrus.WithLevels(codeToLevel),
),
grpc_recovery.UnaryServerInterceptor(
grpc_recovery.WithRecoveryHandlerContext(recoveryHandler),
),
RespHeadersUnaryServerInterceptor(),
grpc_auth.UnaryServerInterceptor(authFunc),
PayloadUnaryServerInterceptor(log, PayloadLoggingDecider),
grpc_validator.UnaryServerInterceptor(),
)),
Unary requests are executed in the following way:
- Function
_GroupService_GetGroup_Handler is called first! It calls
the first interceptor but before that, it creates a handler that
wraps the first middleware and passes to the interceptor chain.
- The first interceptor is:
grpc_ctxtags.UnaryServerInterceptor().
It calls the handler passed, which is the next interceptor.
- The next interceptor is
grpc_logrus.UnaryServerInterceptor and so on.
At some point, we are calling the interceptor executing authentication.
- The last interceptor (
grpc_validator.UnaryServerInterceptor()) calls
finally handler created by GroupService_GetGroup_Handler.
- First middleware is called. The call is executed through the middleware
chain, and may reach the core server, but may return earlier.
- Interceptors are unwrapping in reverse order.
It is visible how this is called from the ChainUnaryServer implementation
if you look.
Streaming calls are a bit different because we start from the interceptors
themselves:
- gRPC Server instance takes function
_GroupService_WatchGroup_Handler
and casts into grpc.StreamHandler type.
- Object
grpc.StreamHandler, which is a handler for our method, is passed
to the interceptor chain. During the chaining process, grpc.StreamHandler
is wrapped with all streaming interceptors, starting from the last.
Therefore, the most internal StreamHandler will be
_GroupService_WatchGroup_Handler.
grpc_ctxtags.StreamServerInterceptor() is the entry point! It then
invokes the next interceptors, and we go further and further, till we
reach _GroupService_WatchGroup_Handler, which is called by the last
stream interceptor, grpc_validator.StreamServerInterceptor().
- Middlewares are executed in the same way as always.
See the ChainStreamServer implementation if you don’t believe it.
In total, this should give an idea of how the server works and what are
the layers.
5.2 - Goten controller library
Understanding the Goten controller library.
You should know about controller design from the
developer guide.
Here we give a small recap of the controller with tips about code paths.
The controller framework is part of the wider Goten framework. It has
annotations + compiler parts, in:
You can read more about the Goten compiler.
For now, in this place, we will talk just about generated controllers.
There are some runtime elements for all controller components (NodeManager,
Node, Processor, Syncer…) in runtime/controller direction in Goten
repo: https://github.com/cloudwan/goten/tree/main/runtime/controller.
In the config.proto, we have node registry access config and nodes manager
configs, which you should already know from controller/db-controller config
proto files.
A bit more interesting thing we have with Node managers. As it was said in
the Developer Guide, we scale horizontally by adding more nodes. To have
more nodes in a single pod, which increases the chance of fairer workload
distribution, we often have more than 1 Node instance per type. We organize
them with Node Managers. You should see a directory
runtime/controller/node_management/manager.go.
Each Node must implement:
type Node interface {
Run(ctx context.Context) error
UpdateShardRange(ctx context.Context, newRange ShardRange)
}
Node Manager component creates on the startup as many Nodes as it has
in the config. Next, it runs all of them, but they don’t get yet any share
of shards. Therefore, they are idle. Managers register all nodes in
the registry, where all node IDs across all pods are collected.
The registry is responsible for returning the shard range assigned for
each node. Whenever a pod dies or a new one is deployed, the Node
registry will notify the manager about new shard ranges per Node. It then
notifies the relevant Node via the UpdateShardRange call.
Registry for Redis uses periodic polling, therefore there may be a chance
two controllers executing the same work in theory for a couple of seconds.
It probably will be better to improve, but we design controllers around
the observed/desired state, and duplicating the same request may bring some
temporary warning errors, but they should be harmless. Still, it’s a field
for improvement.
See the NodeRegistry component (in file registry.go, we use Redis).
Apart from the node managers directory in runtime/controller, you can see
the processor package. We have there from more notable elements:
- Runner module, which is processor runner goroutine. It is the component
for executing all events in a thread-safe manner, but developers must not
do any IO.
- Syncer module, which is generic and based on interfaces, although we
generate type-safe wrappers in all controllers. It is quite large, it
consists of Desired/Observed state objects (file
syncer_states.go),
an updater that operates on its own goroutine (file syncer_updater.go),
and finally central Syncer object, defined in syncer.go. It compares
the desired vs observed state and pushes updates to the syncer updater.
- In
synchronizable we have structures responsible for propagating
sync/lostSync events across Processor modules, so ideally developers
don’t need to handle them themselves.
Syncer is fairly complex, it needs to handle failures/recoveries, resets,
and bursts of updates. Note that it does not use Go channels because:
- They have limited capacity (defined). This is not nice considering we
have IO works there.
- Maps are best if there are multiple updates to a single resource
because they will allow to merging of multiple events (overwrite
previous ones). Channels would force at least to consume all items
from the queue.
5.3 - Goten data store library
Understanding the Goten data store library.
The developer guide gives some examples of simple interaction with
the Store interface, but hides all implementation details, which we
will cover here now, at least partially.
The store should provide:
- Read and write access to resources according to the resource Access
interface. Transactions, which will guarantee resources that have been
read from the database (or query collections) will not change before
the transaction is committed. This is provided by the core store module,
described in this doc.
- Transparent cache layer, reducing pressure on the database, managed
by “cache” middleware, described in this doc.
- Transparent constraint layer handling references to other resources,
and handling blocking references. This is a more complex topic, and
we will discuss this in different documents (multi-region, multi-service,
multi-version design).
- Automatic resource sharding by various criteria, managed by store
plugins, covered in this doc.
- Automatic resource metadata updates (generation, update time…),
managed by store plugins, covered in this doc.
- Observability is provided automatically (we will come back to it
in the Observability document).
The above list should however at least give an idea, that interface calls
may be often complex and require interactions with various components using
IO operations! In general, a call to the Store interface may involve:
- Calling underlying database (mongo, firestore…), for transactions
(write set), non-cached reads…
- Calling cache layer (redis), for reads or invalidation purposes.
- Calling other services or regions in case of references to resources
to other services and regions. This will be not covered by this document
but in this multi-multi-multi thing.
Store implementation resides in Goten, here:
https://github.com/cloudwan/goten/tree/main/runtime/store.
The primary file is store.go, with the following interfaces:
Store is the public store interface for developers.
Backend and TxSession are to be implemented by specific backend
implementations like Firestore and Mongo. They are not exposed to
end-service developers.
SearchBackend is like Backend, for just for search, which is often
provided separately. Example: Algolia, but in the future we may introduce
Mongo combining both search and regular backend implementation.
The store is also actually a “middleware” chain like a server. In the file
store.go we have store struct type, which wraps the backend and
provides the first core implementation of the Store interface. This wrapper
does:
- Add tracing spans for all operations
- For transactions, store an observability tracker in the current
ctx object.
- Invokes all relevant store plugin functions, so custom code can be
injected apart from “middlewares”.
- Accumulates resources to save/delete, does not trigger updates
immediately. They are executed at the end of the transaction.
You can consider it equivalent to a server core module (in the middleware
chain).
To study the store, you should at least check the implementation of
WithStoreHandleOpts.
- You can see that plugins are notified about new and finished
transactions.
- Function runCore is a RETRY-ABLE function that may be invoked again
for the aborted transaction. However, this can happen only for SNAPSHOT
transactions. This also implies that all logic within a transaction must
be repeatable.
- runCore executes a function passed to the transaction. In terms of
server middleware chains, it means we are executing outer + custom
middleware (if present) and/or core server.
- Store plugins are notified when a transaction is attempted (perhaps again),
and get a chance to inject logic just before committing. They also have
a chance to cancel the entire operation.
- You should also note, that Store Save/Delete implementations do not add
any changes to the backend. Instead, creations, updates, and deletions
are accumulated and passed in batch commit inside
WithStoreHandleOpts.
Notable things for Save/Delete implementations:
- They don’t do any changes yet, they are just added to the change set
to be applied (inside WithStoreHandleOpts).
- For Save, we extract current resources from the database and this is
how we detect whether it is an update or creation.
- For Delete, we also get the current object state, so we know the full
resource body we are about to delete.
- Store plugins get a chance to see created/updated/deleted resource
bodies. For updates, we can see before/after.
To see a plugin interface, check the plugin.go file. Some simple store
plugins you could check, are those in the directory store_plugins:
- metaStorePlugin in
meta.go must be always the first store plugin
inserted. It ensures the metadata object is initialized and tracks
the last update.
- You should also see a sharding plugins (
by_name_sharding.go and
by_service_id_sharding.go),
Multi-region plugins and design will be discussed in another document.
Store Cache middleware
The core store module, as described in store.go, is wrapped with cache
“middleware”, see subdirectory cache, file cached_store.go, which
implements the Store interface and wraps the lower level:
- WithStoreHandleOpts decorates function passed to it, to include cache
session restart, in case we have writes that invalidate the cache.
After WithStoreHandleOpts finishes (inner), we need to push invalidated
objects to the worker. It will either invalidate or mark itself as bad
if invalidation fails.
- All read requests (Get, BatchGet, Query, Search) first try to get data
from the cache and pass it to the inner in case of failure, cache miss,
or not cache-able.
- Struct
cachedStore implements not only the Store interface but the
store plugin as well. In the constructor NewCachedStore you should
see it adds itself as a plugin. The reason is that cachedStore is
interested in creating/updated (pre + post) and deleted resource
bodies. Save provides only the current resource body, and Delete
provides only the name to delete. To utilize the fact that the core
store already extracts the “previous” resource state, we implement
cachedStore as a plugin.
Note that watches are non-cacheable. The cached store also needs a separate
backend, we support as of now Redis implementation only.
The reason why we invalidate references/query groups after the transaction
concludes (WithStoreHandleOpts), is because we want new changes to be
already in the database. If we invalidate after writes, then when the new
cache is refreshed, it will be for data after the transaction. This is one
safeguard, but not sufficient yet.
The cache is written to during non-transaction reads (gets or queries). If
results were not in the cache, we fall back to the internal store, using
the main database. With results obtained, we are saving them in cache, but
this is a bit less simple:
- When we first try to READ from cache but face cache MISS, then we are
writing “reservation indicator” for the given cache key.
- When we get results from an actual database, we have fresh results…
but there is a small chance, there is a write transaction undergoing,
that just finished and invalidated cache (deleted keys).
- Cache backend writer must update cache only if data was not invalidated,
if reservation indicator was not deleted, then no write transaction
happened. We can safely update the cache.
This reservation is not done in cached_store.go, it is required behavior
from the backend, see store/cache/redis/redis.go file. It uses SETXX when
updating the cache, meaning we write only if data exists (reservation marker
is present). This behavior is the second safeguard for a valid cache.
The remaining issue may potentially be with 2 reads and one writing
transaction:
- First read request faces, cache miss, makes reservation.
- First read request gets old data from the database.
- Transaction just concluded, overwriting old data, deleting reservation.
- Second read also faces cache miss, and makes a reservation.
- The second read gets new data from the database.
- Second read updates cache with new data.
- First request updates cache with old data, because key exists
(redis only supports if key exists condition)!
This is a known scenario that can cause the issue, it however relies on
the first read request being suspended for quite a long time, allowing
for concluded transaction, and invalidation (which happens with extra
delay after write), furthermore we have full flow of another read
request. As of now, probability may be comparable to serial accidental
lotto wins, so we still allow for the long-live cache. Cache update
happens in the code just after getting results from the database, so
first read flow must be suspended by the CPU scheduler for quite
a very long and then starved a bit.
It may have been better if we find a Redis alternative, that can do
proper Compare and Swap, cache update can only happen for reservation
key, and this key must be unique across read requests. It means the
first request will be only written if the cache contains the reservation
key with the proper unique ID relevant to the first request. If it
contains full data or the wrong ID, it means another read updates
reservation. If some read has cache miss, but sees a reservation mark,
then it must skip cache updating.
The cached store relies on the ResourceCacheImplementation interface,
which is implemented by code generation, see any
<service>/store/<version>/<resource> directory, there is a cache
implementation in a dedicated file, generated based on cache annotations
passed in a resource.
Using centralized cache (redis) we can support very long caches, lasting
even days.
Each resource has a metadata object, as defined in
https://github.com/cloudwan/goten/blob/main/types/meta.proto.
The following fields are managed by store modules:
create_time, update_time and delete_time. Two of these are
updated by the Meta store plugin, delete is a bit special since
we don’t have yet a soft delete function, we have asynchronous
deletion and this is handled by the constraint store layer,
not covered by this document.
resource_version is updated by Meta store plugin.
shards are updated by various store plugins, but can accept
client sharding too (as long as they don’t clash).
syncing is provided by a store plugin, it will be described
in multi-region, multi-service, multi-version design doc.
lifecycle is managed by a constraint layer, again, it will be
described in multi-region, multi-service, multi-version design doc.
Users can manage exclusively: tags, labels, annotations, and
owner_references, although the last one may be managed by services
when creating lower-level resources for themselves.
Field services is often a mix: Each resource may often apply its
own rules. Meta service populates this field itself, For IAM, it
depends on kind: For example, Roles and RoleBindings detect their
contents and decide what services own them and which can read them.
When 3rd party service creates some resource in core SPEKTRA Edge, they
must annotate their service. Some resources, like Device in
devices.edgelq.com, its the client deciding which services
can read it.
Field generation is almost dead, as well as uuid. We may however
fix this at some point. Originally Meta was copied and pasted from
Kubernetes and not all the fields were implemented.
Auxiliary search functionality
The store can provide Search functionality if this is configured. By
default, FailedPrecondition will be returned if no search backend exists.
As of now, the only backend we support is Algolia, but we may add Mongo
as well in the future.
If you check the implementation of Search in store.go and
cache/cached_store.go, it is pretty much like List, but allows additional
search phrases.
Since the search database is however additional to the main one, there
is some problem to resolve: Syncing from the main database to search.
This is an asynchronous process, and the Search query after Save/Delete
is not guaranteed to be accurate. Algolia says it may even be minutes
in some cases. Plus, this synchronization must not be allowed within
transactions, because there is a chance search backend can accept updates,
but the primary database not.
The design decisions regarding search:
- Updates to the search backend are happening asynchronously after
the Store’s successful transaction.
- Search backend needs separate cache keys (they are prefixed), to
avoid mixing.
- Updates to the search backend must be retried in case of failures
because we cannot allow the search to stay out of sync for too long.
- Because of potentially long search updates and, the asynchronous nature
of them, we decided that search writes are NOT executed by Store
components at all! The store does only search queries.
- We dedicated a separate
SearchUpdater interface (See
store/search_updater.go file) for updating the Search backend.
It is not a part of the Store!
- The
SearchUpdater module is used by db-controllers, which observe
changes on the Store in real-time, and update the search backend
accordingly, taking into account potential failures, writes must
be retried.
- Cache for search backend needs invalidation too. Therefore, there
is a
store/cache/search_updater.go file too, which wraps the inner
SearchUpdater for the specific backend.
- To summarize: Store (used by Server modules) makes Search queries,
DbController using SearchUpdater makes writes and invalidates search
cache.
Other store interface useful wrappers
To achieve a read-only database entirely, use the NewReadOnlyStore
wrapper in with_read_only.go.
Normally, the store interface will reject even reads when no transaction
was set (WithStoreHandleOpts was not used). This is to prevent people from
using DB after forgetting to set transactions explicitly. It can be
corrected by using the WithAutomaticReadOnlyTx wrapper in the
auto_read_tx_store.go.
To also be able to write to a database without transaction set explicitly
using WithStoreHandleOpts, it is possible to use WithAutomaticTx wrapper
in auto_tx_store.go, but it is advised to consider other approaches first.
Db configuration and store handle construction
Store handle construction and database configuration are separated.
The store needs configuration because:
- Collections may need pre-initialization.
- Store indices may need configuration too.
Configuration tasks are configured by db-controller runtimes by convention.
Typically, in main.go files we have something like:
senvstore.ConfigureStore(
ctx,
serverEnvCfg,
v1Desc.GetVersion(),
v1Desc,
schemaclient.GetSchemaMixinDescriptor(),
v1limmixinclient.GetLimitsMixinDescriptor(),
)
senvstore.ConfigureSearch(ctx, serverEnvCfg, v1Desc)
The store is configured after being given the main service descriptor,
plus all the mixins, so they can configure additional collections.
If a search feature is used, then it needs a separate configuration.
Configuration functions are in the
edgelq/common/serverenv/store/configurator.go file, and they refer
to further files in goten:
goten/runtime/store/db_configurator.go
goten/runtime/store/search_configurator.go
Configuration therefore happens at db-controller startup but in
a separate manner.
Then, the store handler we construct in the server and db-controller
runtimes. It is done by the builder from the edgelq repository, see
the edgelq/common/serverenv/store/builder.go file. If you have seen
any server initialization file (main.go), you can see how
the store builder constructs “middlewares” (WithCacheLayer,
WithConstraintLayer), and adds plugins executing various functions.
6 - Goten as a runtime
Understanding the runtime aspect of the Goten framework.
Directory runtime contains various libraries linked during compilation.
Many more complex cases will be discussed throughout this guide, here is
rather a quick recap of some common/simpler ones.
runtime/goten
It is rather tiny, and mostly defines interface GotenMessage, which just
merges fmt.Stringer and proto.Message interfaces. Any message generated
by protoc-gen-goten-go implements this interface. We could use it to
figure out who generated the interface.
runtime/object
For resources and many objects, but excluding requests/responses, Goten
generates additional helper types. This directory contains interfaces
for them. Also, for each proto message that has those helper types, Goten
generates implementation as described in the interface GotenObjectExt.
-
FieldPath
Describes some path valid within the associated object.
-
FieldMask
Set of FieldPath objects, all valid for the same object.
-
FieldPathValue
Combination of FieldPath and valid underlying value.
-
FieldPathArrayOfValues
Combination of FieldPath and valid list of underlying values.
-
FieldPathArrayItemValue
Combination of FieldPath describing slice and a valid underlying
item value.
runtime/resource
This directory Contains multiple interfaces related to resource objects.
The most important interface is Resource, which is implemented by every
proto message with Goten resource annotation, see file resource.go. The
next most important probably is Descriptor, as defined in the
descriptor.go file. You can access proto descriptor using
ProtoReflect().Descriptor() call on any proto message, this descriptor
contains additional functionality for resources.
Then, you have plenty of helper interfaces like Name, Reference, Filter,
OrderBy, Cursor, and PagerQuery.
In the access.go file you have an interface that can be implemented by
a store or API client by using proper wrappers.
Note that resources have a global registry.
runtime/client
It contains important descriptors: For methods, API groups, and the whole
service, but within a version. It has some narrow cases, for example in
observability components, where we get request/response objects, and we
need to use descriptors to get something useful.
More often we use service descriptors, mostly for convenience for finding
methods or more often, iterating resource descriptors.
It contains a global registry for these descriptors.
runtime/access
This Directory is connected with access packages in generated services,
but it is relatively poor because those packages are pretty much
code-generated. It has mostly interfaces for watcher-related components.
It has however powerful registry component. If you have a connection to
the service (just grpc.ClientConnInterface) and a descriptor of
the resource, you can construct basic API Access (CRUD) or a high-level
Watcher component (or lower-level QueryWatcher). See the
runtime/access/registry.go file for the actual implementation.
Note that this global registry needs to be populated, though. When a specific
access package — <service>/access/<version>/<resource> — is imported, its
init function calls this global registry and stores the constructors.
This is the reason we have so many “dummy” imports, just to invoke init
functions, so some generic modules can create access objects they need.
runtime/clipb
This contains a set of common functions/types used by CLI tools, like
cuttle.
runtime/utils
This directory is worth mentioning for its proto utility functions, like:
-
GetFieldTypeForOneOf
From the given proto message, can be empty, dummy, extract the actual
reflection type under the specified oneof paths. Not an interface, but
the final path. Normally it takes some effort to get it…
-
GetValueFromProtoPath
From given proto object and path, extracts single current value. It
takes into account all Goten specific types, including in oneofs. If
the last item is an array, it returns the array as a single object.
-
GetValuesFromProtoPath
Like GetValueFromProtoPath, but returns multiple values, if
the field path points to a single object, it is a one-element array.
If the field path points to some array, then it contains an array of
those values. If the last field path item is NOT an array, but some
middle field path item is an array, it will return all values, making
this more powerful than GetValueFromProtoPath.
-
SetFieldPathValueToProtoMsg
It sets value to a proto message under a given path. It allocates all
the paths in the middle if sub-objects are missing, and resets oneofs
on the path.
-
SetFieldValueToProtoMsg
It sets a value to a specified field by the descriptor.
In this section
6.1 - Observability
Understanding the observability module in the Goten runtime.
In the Goten repo, there is an observability module located at
runtime/observability. This module is for:
- Store tracing spans (Jaeger and Google Tracing supported)
- Audit (Service
audit.edgelq.com).
- Monitoring usage (Metrics are stored in
monitoring.edgelq.com).
In the Goten repo, this module is rather small, in observer.go we have
Observer for spans. Goten also stores in context for the current gRPC call
object called CallTracker (call_tracker.go). This generic tracker is
used by the Audit and Monitoring usage reporter.
Goten also provides a global registry, where listeners can tap in to
monitor all calls.
The mentioned module is however just a small base, more proper code is
in SPEKTRA Edge repository, directory common/serverenv/observability:
-
InitCloudTracing
It initializes span tracing. It registers a global
instance, but is picked when we register the proper module, in file
common/serverenv/grpc/server.go. See function NewGrpcServer,
option grpc.StatsHandler. This is where tracing is added to
the server. Audit and Monitoring have not yet been migrated to this
mechanism.
-
InitServerUsageReporter
It initializes usage tracking. First, it stores a global usage
reporter, that periodically sends usage time series data. It also
stores standard observers, usage for store and API. They are registered
in Goten observability module, to catch all calls.
-
InitAuditing
It creates a logs exporter, as defined in the audit/logs_exporter
module, then registers within the goten observability module.
Usage tracking
In the file common/serverenv/observability/observability.go, inside
function InitServerUsageReporter, we initialize two modules:
-
Usage reporter
It is a periodic job, that checks all recorders from time to time,
and exports usage as time series. It’s defined in
common/serverenv/usage/reporter.go.
-
usageCallObserver object
With the RegisterStdUsageReporters call, it is registered within
Goten observability (gotenobservability.RegisterCallObserver)
(defined in common/serverenv/usage/std_recorders.
Reporter is supposed to be generic, there is a possibility to add more
recorders. In the std_recorders directory, we just add a standard observer
for API calls, we track usage on the API Server AND local store usage.
If you look at Reporter implementation, note that we are using always
the same Project ID. This is Service Project ID, global for the whole
Service, shared across all Deployments for this service. Each Service
maintains usage metrics in its project. By convention, if we want to
distinguish usage across user projects, we have a label for it,
user_project_id. This is a common convention. See files
std_recorders/api_recorder.go and std_recorders/storage_recorder.go,
find RetrieveResults calls. We are providing a user_project_id label
for all time series.
Let’s describe standard usage trackers. For this, the central point is
usageCallObserver, defined in the call_observer.go file. If you look
at it, it catches all unnecessary requests/responses plus streams
(new/closed streams, new client or server messages). Its responsibilities
are:
- Insert store usage tracker in the context (via CallTracker).
- Extract usage project IDs from requests or responses (where possible).
- Notify API and storage recorders when necessary, storage recorder
needs periodic flushing for streaming especially.
To track actual store usage, there is a dedicated store plugin, SPEKTRA Edge
repository, file common/store_plugins/usage_observer.go. It gets
a store usage tracker and increments values when necessary!
In summary, this implementation serves to provide metrics for fixtures
defined in monitoring/fixtures/v4/per_service_metric_descriptor.yaml.
Audit
The audit is initialized in general in the
common/serverenv/observability/observability.go file, inside the function
InitAuditing. It calls NewLogsExporter from the
audit/logs_exporter package.
Then, inside RegisterExporter, defined in file
common/serverenv/auditing/exporter.go, we are hooking up two objects
into Goten observability modules:
-
auditMsgVersioningObserver
It’s responsible for catching all request/response versioning
transformations.
-
auditCallObserver
It’s responsible for catching all unary and streaming calls.
Of course, tracking API and versioning is not enough, we also need to
export ResourceChangeLogs somehow. For this, we have also an additional
store plugin in the file common/store_plugins/audit_observer.go file!
It tracks changes happening in the store and pings Exporter when necessary.
When the transaction is about to be committed, we call
MarkTransactionAsReady. It may look a bit innocent, but it is not,
see implementation. We are calling OnPreCommit, which is creating
ResourceChangeLog resources! If we do not succeed, then we return
an error, it will break the entire transaction in result. This is to
ensure that ResourceChangeLogs are always present, even if we fail
to commit ActivityLogs later on, so something is still there in audit.
The reason why we have a separate common/serverenv/auditing directory
from the audit service, was some kind of idea that we should have
an interface in the “common” part, but implementation should be elsewhere.
This was an unnecessary abstraction, especially since we don’t expect
other exporters here (and we want to maintain functionality and be able
to break it). But for now, it is still there and probably will stay due
to low harm.
Implementation of the audit log exporter should be fairly simple, see
the audit/logs_exporter/exporter.go file in the SPEKTRA Edge repository.
Basically:
-
IsStreamReqAuditable and IsUnaryReqAuditable are used
to determine whether we want to track this call. If not, no further
calls will be made.
-
OnPreCommit and OnCommitResult are called to send
ResourceChangeLog. Those are synchronous calls, they don’t exist
until the Audit finishes processing. Note that it will extend a bit
duration of store transactions!
-
OnUnaryReqStarted and OnUnaryReqFinished are called for unary
requests and responses.
-
OnRequestVersioning and OnResponseVersioning are called for
unary requests when their bodies are transformed between API versions.
The function of it is to extract potential labels from updated
requests or responses. Activity logs recorded will still be done
for the old version.
-
OnStreamStarted and OnStreamFinished should be self-explanatory.
-
OnStreamExportable notifies when ActivityLog can be generated.
It is used to send ActivityLogs before the call finishes.
-
OnStreamClientMessage and OnStreamServerMessage add client/server
messages to ActivityLogs.
-
OnStreamClientMsgVersioning and OnStreamServerMsgVersioning
notify the exporter when client or server messages are transformed
to different API versions.
Notable elements:
- Audit log exporter can sample unary requests when deciding whether
to audit or not.
- While ResourceChangeLog is sent synchronously and extends call duration,
ActivityLogs does not. The audit exporter maintains a set of workers
for streaming and unary calls, they have a bit of a different
implementation. They work asynchronously.
- Stream and unary log workers will try to accumulate a small batch
of activity logs before sending, them to save on IO work. They have
timeouts based on log size and time.
- Stream and unary log workers will retry failed logs, but if they
accumulate too much, they will start dropping.
- Unary log workers send ActivityLogs for finished calls only.
- Streaming log workers can send ActivityLogs for ongoing calls.
If this happens, many Activity log fields like labels are no
longer updateable. But request/responses and exit codes will be
appended as Activity Log Events.