v6.4.7
Transparent HTTP proxy timeouts, tighter workload security contexts
Notable new features⚓︎
First-class Organizations (multi-org)⚓︎
A single Hydrolix installation can now serve multiple customer organizations, each with its own configuration, projects, and query options. The intake services and query engine load per-organization configuration, and the Hydrolix UI adds a customer switcher. The Config API's complete set of v2 endpoints supports the multi-org model with shortened URLs that take the organization from the request body instead of a path parameter. See the Organizations guides to plan and enable the transition. See breaking changes.
Tighter security contexts for cluster workloads⚓︎
Operator-managed workloads now run under tightened Kubernetes security contexts by default. This includes reduced capabilities across pods and a strict read-only container root filesystem where supported. Services that need space for file CRUD (for example, gunicorn control sockets, stream_head zip decompression, and Grafana plugin directories) receive dedicated writable emptyDir and /tmp mounts, so ensure this security hardening is transparent to running workloads. A small number of services that must run as root keep an explicit override.
Faster partition cleanup⚓︎
Request counts to object stores have been greatly reduced during partition cleanup, making it easier to stay under provider rate limits. The reaper depends on each object store's bulk API and backs off when rate-limited. Also, the decay and merge-cleanup services sweep tables in interleaved batches so one backlogged table can't delay the rest.
Breaking changes⚓︎
- Project to customer assignment is immutable. The
PUTandPATCH /projects/:idendpoints return an HTTP 400 if thecustomerparameter in the request body differs from the existing object. - Customer ID is always required on projects now.
POST /projectsrequires a customer ID. See also Example changed workflows. - Deployment ID moves from project to customer. Accounts with new permission
deployment_id_customercan use thePUT /customers/:id/deployment_idto change deployment IDs. It completely replacesPUT /projects/:id/deployment_id. Projects continue to return the deployment ID as a read-only field derived from the customer. Accounts which could previously modify deployment IDs have the new permission. - Credentials and storages can't be detached from a customer if in use. The
POST /customers/:id/remove_credentialandPOST /customers/:id/remove_storageendpoints return HTTP 400 if the credential or storage is still in use by any tables belonging to the customer. - Storages must be assigned to a customer before use. Tables can't use a storage object until it's assigned to the table's customer. Use
POST /customers/:id/add_storage. - Credentials must be assigned to a customer before use. Tables, sources, and jobs can't use a credential until it's assigned to the customer. Use
POST /customers/:id/add_credential. Credentials must belong to every customer using the storage to be used with the storage.
Upgrade instructions⚓︎
Upgrade steps for this release live on their own page: Upgrade to v6.4.x. Start there once you've read the breaking changes.
Changelog⚓︎
Organizations and tenancy⚓︎
Multi-org publishing is on by default as of this release. The Config API uses a new v3 storage format that supports configuration for multiple organizations in a cluster. The query and intake services consume it with no tunables to set, so every installation moves onto the new config path, not only the ones running more than one organization. That config format version is unrelated to the Config API endpoint version, v2, which the Configuration and APIs section covers. The rest is the plumbing that made it possible, plus customer management in the web UI. Project ownership also becomes required and frozen once set - both covered under breaking changes, and the freeze needs action before you upgrade.
Changes⚓︎
-
Multi-org (v3) configuration publishing is now enabled by default.
turbine_api_feature_publish_multi_org_configs,turbine_core_enable_multi_org, andintake_use_multi_org_configall default totrue, so the Config API publishes per-organization config and the query and intake services consume it without extra configuration. -
Required a customer relationship for every new or modified project. Moved the
hdx_deployment_idattribute, formerly on the project to the customer and granted a new RBAC permissiondeployment_id_customerto any accounts which helddeployment_id_project. This also stops and deletes any background jobs started in Hydrolix v6.1 that created customers from projects. -
Added a background job to create customer-level query options and merge pool settings from the original, single organization object in the cluster. Project-level query options and merge pools are untouched.
-
The query engine can now load per-organization (v3) configuration, the core piece of the first-class organizations feature. Each org's config syncs and reports its revision independently. An org can carry
default_query_options, resolved per query with org < project < table precedence. The Organizations guides cover enabling it. -
The
turbine_core_enable_multi_orgtunable now defaults totrue, so the query engine is ready to consume the v3 (multi-org) config format out of the box. Multi-org isn't active end to end untilturbine_api_feature_publish_multi_org_configsis also enabled. -
Intake services are compatible with per-organization configuration (v3), part of the first-class organizations feature. Enable it with the new
intake_use_multi_org_configtunable (defaultfalse) together withturbine_api_feature_publish_multi_org_configs. -
Changes related to the first-class organizations update. Updates the configuration loader to support the v3 config format. V3 requires storages and credentials be taken from a global config file rather than individual org files.
-
Updated the Rust config loader to read v3 of the config format, which stores each org in its own file. Org configs are now cached by revision number, so unchanged configs aren't reparsed.
-
Updated the
/config-versionendpoint for ingest services to handle both v3 and v2 config.json formats. When v3 format is present,config_customer_revisionis used. Otherwise, the prior v2 behavior is unchanged. -
Updated the spread list storage definition to understand the v3 config.json format. The v2 config.json format held the table and storage definitions in a single file, where the v3 format manages them independently, the storage definitions in the global file and the table definitions in a per organization file.
-
Added upgrade migration logic to retain database columns rather than immediately remove the field and the column. By retaining the
deployment_idin the database, administrators can safely upgrade and downgrade again without incurring any data loss. A future release can drop the vestigial column. -
On multi-org (v3) clusters, the query engine's
/config-versionendpoint now reports the customer (v3) revision, matching the intake services. -
The version-service now accepts an org ID and queries pods at
/org-config-versionfor per-organization config versions when multi-org is enabled. It also skips pods whose version response can't be parsed, instead of recording them as version 0 and dragging the reported lowest version down. -
The web UI supports multi-org: a header dropdown switches between the customers a user can access, and the Tables page lists the selected customer's projects. Project creation, dictionaries, functions, and user listing remain org-based.
-
The web UI adds customer management: create, edit, and delete customers, and add or remove a customer's users. This is part of the first-class organizations work.
Identity and access⚓︎
Authorization and audit behavior. Introduced a new Config API requirement: a user must share a customer with any object they operate on, so automation that reaches across customers starts getting rejected. The fixes correct authorization decisions and audit cleanup that were wrong in specific configurations, and land without anything to configure. The tightened workload security contexts are covered under notable new features.
Changes⚓︎
-
The Config API now requires a user to share a customer with any object they operate on. Requests for objects outside the user's customers are rejected.
-
Identify user by UUID rather than email in Traefik auth logs.
-
Queries authenticated with a service account token now record the caller as
<service account ID>/<token ID>in theusercolumn ofhydro.logsand inhdx.active_queries. Previously the user was blank, and the raw token could be substituted as the user name. Username and password logins and bearer tokens that carry a username are unchanged.
Fixes⚓︎
-
Fixed the Traefik auth plugin failing to refresh user permissions for requests that authenticate with Basic auth, because the permissions endpoint accepts only Bearer tokens. The plugin now obtains an access token first, and no longer shares auth state between concurrent requests.
-
Fixed the Traefik auth plugin applying RBAC authorization to Traefik users, which have no RBAC of their own. Authorization is now skipped for Traefik users and still enforced for turbine-api users.
-
Fixed the daily Keycloak audit-cleanup task so migrated Keycloak records are purged and its data-loss safeguard engages as intended.
-
Fixed org- and project-scoped permissions on the v2 endpoints so creating a project or an org- or project-level object is correctly scoped to its parent. Table-scoped creation already worked.
Cluster operations⚓︎
Two defaults changed: the default HDX Scaler (hdx-scaler -> hdx-scaler-go), and the priority level of query pods according to the Kubernetes scheduler. Read the first two entries before upgrading a cluster with custom scaler configuration. The rest hardens Traefik and the operator against failure modes.
Changes⚓︎
-
hdx-scaler-gois now the default autoscaling engine and runs on every cluster. On upgrade,hdxscalersblocks without anenginekey move from the Python scaler tohdx-scaler-go. -
Query deployments now use the
hdx-highestpriority class by default, matching the intake components. This makes the Kubernetes scheduler less likely to evict or preempt query pods under resource pressure. -
Added readiness and liveness endpoints to the HTTP proxy (
chproxy), along with areadiness_stuck_timeoutserver tunable (default30s) that controls how long the proxy stays in the readiness-to-liveness escalation state before it's considered unhealthy. -
The operator-resources install URL now accepts an
image-pull-secretquery parameter that sets a Kubernetes image pull secret on the operator. Use it when operator images are hosted in a private registry that requires authentication. -
Add optimization logic to both HDX Scalers to account for 0 -> 1 scaling, such as in the case of periodic tasks such as batch ingest. Previously, the HDX Scalers were primarily optimized for the scaling from 1 -> N case.
Fixes⚓︎
-
Fixed looping operator reconciliation failure. Earlier, a pool specifying a replica range without any defined hdxscalers triggered a reconciliation failure. Now, the default horizontal pod autoscaler is used, preventing the stuck operator and recurring reconciliation failure.
-
hdx-scaler-go now re-registers its admission webhook's certificate with Kubernetes every five minutes, so a rotated certificate is trusted without a pod restart. Previously it registered the certificate only at startup.
-
Corrected Traefik HTTP pool routing selection logic to match an exact pool path, like
/pool/http-head-1or an anchored path prefix/pool/http-head-1/. Earlier, the reverse proxy could misroute/pool/http-head-10by matching the unanchored prefix/pool-http-head-1. -
Fixed inconsistency in operator calculation of Traefik grace period timeout. Earlier, the grace period wasn't guaranteed to cover the container preStop hook and maximum keep alive time, which caused broken client connections during Traefik scaling events. Now, the calculation accounts for container preStop and
traefik_keep_alive_max_timedurations. -
Corrected scale lookup for query head Config API sidecar which provides authorization, permissions, and data access control information. Earlier, the sidecar inherited unnecessarily large resource requirements of the query head.
-
Fixed the
silence-linodeDaemonSet failing to start on clusters withsilence_linode_alertsenabled. The tighter security contexts in this release made its container root filesystem read-only, and it had no writable temp directory, so nodes added after upgrading kept Linode's default alert thresholds. Fixed by adding a writableemptyDirvolume at/tmp. -
Fixed merge controller, spill controller, and validator pods crash looping during upgrade. They started before the new
init-turbine-apiandinit-clusterjobs had published their configuration, because theDONEstatus left by the previous version's init jobs let them through. The init jobs now record which release finished, and pods wait for the current release.
Query⚓︎
Query no longer enforces a deadline by default. The HTTP proxy used to cut off long-running queries; both of its limits now default to no limit, so any ceiling has to be set manually either at query head or by overriding the proxy defaults for the cluster.
This release also reduced how much of the catalog a query reads, and renames a #.catalog column (billing_byte -> billing_bytes), so any saved query or dashboard referencing the old name needs updating.
Changes⚓︎
-
Updated the HTTP proxy (
chproxy) default configuration to ensure it operates transparently by removing its limits, leaving enforcement to the query-head. Previously, the proxy killed long-running queries. Both proxy-enforced limits now default to no limit, so a query runs as long as the query head allows. Explicit per-cluster overrides still work if you want the proxy to enforce a deadline. Also upgraded chproxy tov0.6.8which treats an omitted timeout as no limit.http_proxy.users.max_execution_time:2m->0shttp_proxy.server.write_timeout:4m->0s
-
Updated the HTTP proxy (
chproxy) default version tov0.6.8. -
Summary tables now prune partitions by shard key, so a query with a shard-key filter no longer scans the whole catalog.
-
Partition Bloom filters whose false-positive rate would exceed the new
hdx_partition_bloom_max_fpr_pctthreshold are now skipped instead of built at full size for little benefit, trimming manifest size on high-cardinality columns. -
Hydrolix deploys a HDX Query Discovery service sidecar to each
query-headandturbine-apipod whendisable_zookeeper: trueorhdx_query_discovery_service.enable: true. -
Time filters that wrap a
DateTimeprimary key intoDateTime,toUnixTimestamp,toUInt32,toUInt64,toTimeZone, orCASTnow prune partitions like the bare column and satisfyhdx_query_timerange_required. Before, these queries read every partition, or were rejected whenhdx_query_timerange_requiredwas enabled. Other functions, such astoStartOfHour, andDateTime64primary keys keep the previous behavior. Thehdx_query_timerange_requirederror message now lists the accepted conversions.
Fixes⚓︎
-
Renamed the
#.catalogvirtual table'sbilling_bytecolumn tobilling_bytes, matching the name used elsewhere, and fixedWHEREfilters on it that silently returned no rows. Update any query that selected or filteredbilling_byteto usebilling_bytes. -
Prevented the query head from a deadlock race condition. Inverted namespace and storage map lock acquisition order in independent threads, allowed each to block on the other. Bug occurred under configuration reload while listing tables concurrently. Now both operations use the same lock acquisition order.
-
Corrected a flaw in ClickHouse native TCP authentication handling which allowed clients to bypass authentication.
-
Added a tunable to control an unauthenticated Prometheus remote read endpoint on the ClickHouse HTTP service. By default,
enable_prom_readturns off access. The/prom_readendpoint didn't require authentication and allowed SQL queries built from client-supplied parameters. -
Removed query options
hdx_log_queryandhdx_internal_query, which clients could use to bypass query logging or audit records. These query options are now ignored completely. Queries including the options in aSETTINGSclause continue to work. -
Prevented use of network-capable ClickHouse table functions
remote,remoteSecure,mysql,cluster, andclusterAllReplicas. Authenticated users could use these for outbound network connections, a server-side request forgery vulnerability.
Data lifecycle⚓︎
Partition cleanup has two major changes: resiliency to object store rate limits, and faster backlog processing.
Producers to the reaper queue, such as the merge cleanup and decay services, backpressure additional requests when the queue is full. Previously, the reaper queue dropped events under heavy load. Added several new environment variables for tuning cleanup parallelism. One adds a manual release for catalogs with substantial inactive partition backlogs which allows the periodic service to clean up catalog data and dispatch reaper events. Expect changes in metrics related to object store request volume and partition counts.
Changes⚓︎
-
Partition reaping uses far fewer object store requests, now using bulk deletion. The reaper backs off if rate-limited. Failed deletes now drop the catalog record, leaving the files to the partition cleaner. New
reaper_*metrics track throttled, deferred, and orphaned partitions. -
The reaper now deletes a partition's known files directly instead of listing the partition prefix first, removing a
LISTrequest per reaped partition. A reap that finds none of its expected files now warns and incrementsreaper_reap_all_not_found_totalso orphaned keys (for example, after a storagebucket_pathchange) are visible. SetREAPER_LIST_BEFORE_DELETEto restore the previous list-then-delete behavior. -
Improved reaper queue partition deletion efficiency. All tasks rejoin the back of the queue for an available worker preventing slow storage locations from blocking faster sibling storage locations. Faster storage locations claim more turns and delete faster.
-
Clusters with large partition backlogs clean up faster. The decay and merge-cleanup services now sweep tables in shuffled, interleaved batches, and merge-cleanup fetches locked partitions without sorting.
-
Adapted the decay and merge cleanup operations to test and respect RabbitMQ queue depth before enqueuing additional reaper events. By using back pressure to pace the producers, the system avoids dropping reap events on queue overrun. Also, the periodic service now emits metrics every minute instead of only once at the end of each batch run. This backlog observability improvement is necessary for bulk cleaning operations.
-
Improved throughput for collecting catalog entries to reap. Removed a catalog query bottleneck by using PostgreSQL
SKIP LOCKand allowing multiple routines to concurrently collect disjoint rows for reaping from the catalog. Use environment variablesDECAY_TABLE_CONCURRENCYandMERGE_CLEANUP_TABLE_CONCURRENCYto set maximum query parallelization. -
Improved throughput for collecting catalog entries to reap. Added multiple in-process communication channels when sending reaper events to RabbitMQ. This removed a shared lock contention bottleneck, allowing parallel collectors to send reap events. It depends on the back pressure detection to avoid RabbitMQ overrun and reaper event drop.
-
Improved reaper system efficiency by dequeueing and acknowledging batches of reap events from RabbitMQ instead handling events one at a time. This increases queue throughput, decreases RabbitMQ contention, and keeps parallel reapers busy cleaning inactive partitions from storage.
-
Added a "delete on dispatch" mode to the decay and merge cleanup jobs in the periodic service. This mode helps drain a backlog of inactive partitions that the decay and merge cleanup jobs can't keep up with. When enabled, the query that finds inactive partitions also deletes their catalog rows and returns the necessary information for reaper to clean up the corresponding partition data. As a result, merge cleanup skips an additional "unlock update" step and the reaper no longer deletes the catalog rows later. If a reaper event is lost after its catalog row is deleted, the storage files are orphaned and partition cleaner reclaims them. Set
DECAY_DELETE_ON_DISPATCHorMERGE_CLEANUP_DELETE_ON_DISPATCHto enable this mode. Both default tofalse.
Fixes⚓︎
- Correctly propagate the billing bytes statistic into catalog metadata during merge. Before the fix, the merge system stored zero for
billing_bytesfor all merged partitions, though the originally ingested partitions had the correct value. The bug inhibited reporting and analysis, but didn't affected bills, as merged partitions aren't used for billing.
Observability⚓︎
Worth a look if you run dashboards or alerts on Hydrolix internals. Where spans come from changes this release, and a new metric family arrives. Panels pointing at the retired collectors go blank after the upgrade.
Changes⚓︎
-
Administrators can now see whether the
hdx-pg-monitorservice's PostgreSQL monitoring queries are succeeding and how long they take. Newpgmonitor_*metrics report query executions, failures by cause, duration, and other diagnostic information. -
Retired cluster-internal sidecar OpenTelemetry collectors in favor of a single central collector. The benefits of batched span collection aren't worth the increased cluster-wide resource cost.
-
Switched the Tulugaq metrics to transform management tools to use the v2 Config API endpoints.
Configuration and APIs⚓︎
The v2 Config API is feature-complete as of this release, so anything still on v1 can migrate. The v1 endpoints remain unchanged. Three Hydrolix cluster spec keys (s3_endpoint, premium_storage_class, and region) have been deprecated, and a fix ensures functions and UDFs aren't orphaned when a project is deleted.
Changes⚓︎
-
Migrated the last remaining Config API endpoints to v2, completing the v2 endpoint set.
-
Added parallel v2 Config API variants for most remaining non-nested v1 endpoints. Calling conventions don't change for the non-nested endpoints.
-
Added v2 Config API endpoints for column policies and row policies, managed as standalone resources rather than as part of the table definition.
-
Added v2 Config API endpoints for column and customer objects.
-
Added v2 Config API endpoints for the catalog and partitions at shorter, top-level paths that take the organization from the request rather than a path parameter. The equivalent v1 org-scoped endpoints are unchanged.
-
The Config API
/customersendpoints can return arevision, a version number for the customer's configuration that changes whenever the config does, so clients can detect a change and skip reloading unchanged config. It's opt-in via?include_revision=Truebecause computing the revision adds overhead. -
The Config API user response now includes a
customerslist, so a user's accessible customers can be read in a single call. -
The
s3_endpoint,premium_storage_class, andregionspec keys are now deprecated tunables. Usedb_bucket_endpointin place ofs3_endpoint, which remains a legacy fallback. -
Improve the performance of a task that ensures users in a cluster who have a global permission are associated with all customer objects.
Fixes⚓︎
-
Deleting a project now also removes its functions and UDFs, which were previously left orphaned.
-
Fixed token issuance on the v2 service account endpoint, which was rejecting requests that didn't also supply
nameandorg. -
Corrected several
hydrolix_oteltunable variables. Fixed incorrect descriptions ofhdx_tableandhdx_transformand updated thegrpc_portandhttp_portdefaults to 5317 and 5318, respectively.
Interfaces and tools⚓︎
Changes⚓︎
-
Upgraded the operator-bundled Hydrolix Grafana datasource plugin version to
0.11.0. Clusters that haven't pinned theversionfield use this plugin version up on upgrade. See the plugin's change log for full per-version change logs. -
Updated the Kibana Gateway default version from
v2.1.0→v2.2.0. This defaults to the v2 Config API endpoints and supports theconfig_api_versionof either v1 or v2 in thekibana_gateway_configtunable.
Fixes⚓︎
- Fixed incorrect scale-value validation in the SIEM source creation and edit form.
Bundled component versions⚓︎
Third-party components whose bundled or default version changed in this release. Where a version change also changes behavior, the entry in the relevant area covers the detail.
| Component | Version in v6.4.x |
|---|---|
| Hydrolix Grafana datasource plugin | 0.11.0 |
HTTP proxy (chproxy) |
v0.6.8 |
| Kibana Gateway | v2.2.0 |