Skip to content

v6.4.7

Transparent HTTP proxy timeouts, tighter workload security contexts

Notable new features⚓︎

First-class Organizations (multi-org)⚓︎

A single Hydrolix installation can now serve multiple customer organizations, each with its own configuration, projects, and query options. The intake services and query engine load per-organization configuration, and the Hydrolix UI adds a customer switcher. The Config API's complete set of v2 endpoints supports the multi-org model with shortened URLs that take the organization from the request body instead of a path parameter. See the Organizations guides to plan and enable the transition. See breaking changes.

Tighter security contexts for cluster workloads⚓︎

Operator-managed workloads now run under tightened Kubernetes security contexts by default. This includes reduced capabilities across pods and a strict read-only container root filesystem where supported. Services that need space for file CRUD (for example, gunicorn control sockets, stream_head zip decompression, and Grafana plugin directories) receive dedicated writable emptyDir and /tmp mounts, so ensure this security hardening is transparent to running workloads. A small number of services that must run as root keep an explicit override.

Faster partition cleanup⚓︎

Request counts to object stores have been greatly reduced during partition cleanup, making it easier to stay under provider rate limits. The reaper depends on each object store's bulk API and backs off when rate-limited. Also, the decay and merge-cleanup services sweep tables in interleaved batches so one backlogged table can't delay the rest.


Breaking changes⚓︎

  • Project to customer assignment is immutable. The PUT and PATCH /projects/:id endpoints return an HTTP 400 if the customer parameter in the request body differs from the existing object.
  • Customer ID is always required on projects now. POST /projects requires a customer ID. See also Example changed workflows.
  • Deployment ID moves from project to customer. Accounts with new permission deployment_id_customer can use the PUT /customers/:id/deployment_id to change deployment IDs. It completely replaces PUT /projects/:id/deployment_id. Projects continue to return the deployment ID as a read-only field derived from the customer. Accounts which could previously modify deployment IDs have the new permission.
  • Credentials and storages can't be detached from a customer if in use. The POST /customers/:id/remove_credential and POST /customers/:id/remove_storage endpoints return HTTP 400 if the credential or storage is still in use by any tables belonging to the customer.
  • Storages must be assigned to a customer before use. Tables can't use a storage object until it's assigned to the table's customer. Use POST /customers/:id/add_storage.
  • Credentials must be assigned to a customer before use. Tables, sources, and jobs can't use a credential until it's assigned to the customer. Use POST /customers/:id/add_credential. Credentials must belong to every customer using the storage to be used with the storage.

Upgrade instructions⚓︎

Upgrade steps for this release live on their own page: Upgrade to v6.4.x. Start there once you've read the breaking changes.

Changelog⚓︎

Organizations and tenancy⚓︎

Multi-org publishing is on by default as of this release. The Config API uses a new v3 storage format that supports configuration for multiple organizations in a cluster. The query and intake services consume it with no tunables to set, so every installation moves onto the new config path, not only the ones running more than one organization. That config format version is unrelated to the Config API endpoint version, v2, which the Configuration and APIs section covers. The rest is the plumbing that made it possible, plus customer management in the web UI. Project ownership also becomes required and frozen once set - both covered under breaking changes, and the freeze needs action before you upgrade.

Changes⚓︎

  • Multi-org (v3) configuration publishing is now enabled by default. turbine_api_feature_publish_multi_org_configs, turbine_core_enable_multi_org, and intake_use_multi_org_config all default to true, so the Config API publishes per-organization config and the query and intake services consume it without extra configuration.

  • Required a customer relationship for every new or modified project. Moved the hdx_deployment_id attribute, formerly on the project to the customer and granted a new RBAC permission deployment_id_customer to any accounts which held deployment_id_project. This also stops and deletes any background jobs started in Hydrolix v6.1 that created customers from projects.

  • Added a background job to create customer-level query options and merge pool settings from the original, single organization object in the cluster. Project-level query options and merge pools are untouched.

  • The query engine can now load per-organization (v3) configuration, the core piece of the first-class organizations feature. Each org's config syncs and reports its revision independently. An org can carry default_query_options, resolved per query with org < project < table precedence. The Organizations guides cover enabling it.

  • The turbine_core_enable_multi_org tunable now defaults to true, so the query engine is ready to consume the v3 (multi-org) config format out of the box. Multi-org isn't active end to end until turbine_api_feature_publish_multi_org_configs is also enabled.

  • Intake services are compatible with per-organization configuration (v3), part of the first-class organizations feature. Enable it with the new intake_use_multi_org_config tunable (default false) together with turbine_api_feature_publish_multi_org_configs.

  • Changes related to the first-class organizations update. Updates the configuration loader to support the v3 config format. V3 requires storages and credentials be taken from a global config file rather than individual org files.

  • Updated the Rust config loader to read v3 of the config format, which stores each org in its own file. Org configs are now cached by revision number, so unchanged configs aren't reparsed.

  • Updated the /config-version endpoint for ingest services to handle both v3 and v2 config.json formats. When v3 format is present, config_customer_revision is used. Otherwise, the prior v2 behavior is unchanged.

  • Updated the spread list storage definition to understand the v3 config.json format. The v2 config.json format held the table and storage definitions in a single file, where the v3 format manages them independently, the storage definitions in the global file and the table definitions in a per organization file.

  • Added upgrade migration logic to retain database columns rather than immediately remove the field and the column. By retaining the deployment_id in the database, administrators can safely upgrade and downgrade again without incurring any data loss. A future release can drop the vestigial column.

  • On multi-org (v3) clusters, the query engine's /config-version endpoint now reports the customer (v3) revision, matching the intake services.

  • The version-service now accepts an org ID and queries pods at /org-config-version for per-organization config versions when multi-org is enabled. It also skips pods whose version response can't be parsed, instead of recording them as version 0 and dragging the reported lowest version down.

  • The web UI supports multi-org: a header dropdown switches between the customers a user can access, and the Tables page lists the selected customer's projects. Project creation, dictionaries, functions, and user listing remain org-based.

  • The web UI adds customer management: create, edit, and delete customers, and add or remove a customer's users. This is part of the first-class organizations work.

Identity and access⚓︎

Authorization and audit behavior. Introduced a new Config API requirement: a user must share a customer with any object they operate on, so automation that reaches across customers starts getting rejected. The fixes correct authorization decisions and audit cleanup that were wrong in specific configurations, and land without anything to configure. The tightened workload security contexts are covered under notable new features.

Changes⚓︎

  • The Config API now requires a user to share a customer with any object they operate on. Requests for objects outside the user's customers are rejected.

  • Identify user by UUID rather than email in Traefik auth logs.

  • Queries authenticated with a service account token now record the caller as <service account ID>/<token ID> in the user column of hydro.logs and in hdx.active_queries. Previously the user was blank, and the raw token could be substituted as the user name. Username and password logins and bearer tokens that carry a username are unchanged.

Fixes⚓︎

  • Fixed the Traefik auth plugin failing to refresh user permissions for requests that authenticate with Basic auth, because the permissions endpoint accepts only Bearer tokens. The plugin now obtains an access token first, and no longer shares auth state between concurrent requests.

  • Fixed the Traefik auth plugin applying RBAC authorization to Traefik users, which have no RBAC of their own. Authorization is now skipped for Traefik users and still enforced for turbine-api users.

  • Fixed the daily Keycloak audit-cleanup task so migrated Keycloak records are purged and its data-loss safeguard engages as intended.

  • Fixed org- and project-scoped permissions on the v2 endpoints so creating a project or an org- or project-level object is correctly scoped to its parent. Table-scoped creation already worked.

Cluster operations⚓︎

Two defaults changed: the default HDX Scaler (hdx-scaler -> hdx-scaler-go), and the priority level of query pods according to the Kubernetes scheduler. Read the first two entries before upgrading a cluster with custom scaler configuration. The rest hardens Traefik and the operator against failure modes.

Changes⚓︎

  • hdx-scaler-go is now the default autoscaling engine and runs on every cluster. On upgrade, hdxscalers blocks without an engine key move from the Python scaler to hdx-scaler-go.

  • Query deployments now use the hdx-highest priority class by default, matching the intake components. This makes the Kubernetes scheduler less likely to evict or preempt query pods under resource pressure.

  • Added readiness and liveness endpoints to the HTTP proxy (chproxy), along with a readiness_stuck_timeout server tunable (default 30s) that controls how long the proxy stays in the readiness-to-liveness escalation state before it's considered unhealthy.

  • The operator-resources install URL now accepts an image-pull-secret query parameter that sets a Kubernetes image pull secret on the operator. Use it when operator images are hosted in a private registry that requires authentication.

  • Add optimization logic to both HDX Scalers to account for 0 -> 1 scaling, such as in the case of periodic tasks such as batch ingest. Previously, the HDX Scalers were primarily optimized for the scaling from 1 -> N case.

Fixes⚓︎

  • Fixed looping operator reconciliation failure. Earlier, a pool specifying a replica range without any defined hdxscalers triggered a reconciliation failure. Now, the default horizontal pod autoscaler is used, preventing the stuck operator and recurring reconciliation failure.

  • hdx-scaler-go now re-registers its admission webhook's certificate with Kubernetes every five minutes, so a rotated certificate is trusted without a pod restart. Previously it registered the certificate only at startup.

  • Corrected Traefik HTTP pool routing selection logic to match an exact pool path, like /pool/http-head-1 or an anchored path prefix /pool/http-head-1/. Earlier, the reverse proxy could misroute /pool/http-head-10 by matching the unanchored prefix /pool-http-head-1.

  • Fixed inconsistency in operator calculation of Traefik grace period timeout. Earlier, the grace period wasn't guaranteed to cover the container preStop hook and maximum keep alive time, which caused broken client connections during Traefik scaling events. Now, the calculation accounts for container preStop and traefik_keep_alive_max_time durations.

  • Corrected scale lookup for query head Config API sidecar which provides authorization, permissions, and data access control information. Earlier, the sidecar inherited unnecessarily large resource requirements of the query head.

  • Fixed the silence-linode DaemonSet failing to start on clusters with silence_linode_alerts enabled. The tighter security contexts in this release made its container root filesystem read-only, and it had no writable temp directory, so nodes added after upgrading kept Linode's default alert thresholds. Fixed by adding a writable emptyDir volume at /tmp.

  • Fixed merge controller, spill controller, and validator pods crash looping during upgrade. They started before the new init-turbine-api and init-cluster jobs had published their configuration, because the DONE status left by the previous version's init jobs let them through. The init jobs now record which release finished, and pods wait for the current release.

Query⚓︎

Query no longer enforces a deadline by default. The HTTP proxy used to cut off long-running queries; both of its limits now default to no limit, so any ceiling has to be set manually either at query head or by overriding the proxy defaults for the cluster.

This release also reduced how much of the catalog a query reads, and renames a #.catalog column (billing_byte -> billing_bytes), so any saved query or dashboard referencing the old name needs updating.

Changes⚓︎

  • Updated the HTTP proxy (chproxy) default configuration to ensure it operates transparently by removing its limits, leaving enforcement to the query-head. Previously, the proxy killed long-running queries. Both proxy-enforced limits now default to no limit, so a query runs as long as the query head allows. Explicit per-cluster overrides still work if you want the proxy to enforce a deadline. Also upgraded chproxy to v0.6.8 which treats an omitted timeout as no limit.

    • http_proxy.users.max_execution_time: 2m -> 0s
    • http_proxy.server.write_timeout: 4m -> 0s
  • Updated the HTTP proxy (chproxy) default version to v0.6.8.

  • Summary tables now prune partitions by shard key, so a query with a shard-key filter no longer scans the whole catalog.

  • Partition Bloom filters whose false-positive rate would exceed the new hdx_partition_bloom_max_fpr_pct threshold are now skipped instead of built at full size for little benefit, trimming manifest size on high-cardinality columns.

  • Hydrolix deploys a HDX Query Discovery service sidecar to each query-head and turbine-api pod when disable_zookeeper: true or hdx_query_discovery_service.enable: true.

  • Time filters that wrap a DateTime primary key in toDateTime, toUnixTimestamp, toUInt32, toUInt64, toTimeZone, or CAST now prune partitions like the bare column and satisfy hdx_query_timerange_required. Before, these queries read every partition, or were rejected when hdx_query_timerange_required was enabled. Other functions, such as toStartOfHour, and DateTime64 primary keys keep the previous behavior. The hdx_query_timerange_required error message now lists the accepted conversions.

Fixes⚓︎

  • Renamed the #.catalog virtual table's billing_byte column to billing_bytes, matching the name used elsewhere, and fixed WHERE filters on it that silently returned no rows. Update any query that selected or filtered billing_byte to use billing_bytes.

  • Prevented the query head from a deadlock race condition. Inverted namespace and storage map lock acquisition order in independent threads, allowed each to block on the other. Bug occurred under configuration reload while listing tables concurrently. Now both operations use the same lock acquisition order.

  • Corrected a flaw in ClickHouse native TCP authentication handling which allowed clients to bypass authentication.

  • Added a tunable to control an unauthenticated Prometheus remote read endpoint on the ClickHouse HTTP service. By default, enable_prom_read turns off access. The /prom_read endpoint didn't require authentication and allowed SQL queries built from client-supplied parameters.

  • Removed query options hdx_log_query and hdx_internal_query, which clients could use to bypass query logging or audit records. These query options are now ignored completely. Queries including the options in a SETTINGS clause continue to work.

  • Prevented use of network-capable ClickHouse table functions remote, remoteSecure, mysql, cluster, and clusterAllReplicas. Authenticated users could use these for outbound network connections, a server-side request forgery vulnerability.

Data lifecycle⚓︎

Partition cleanup has two major changes: resiliency to object store rate limits, and faster backlog processing.

Producers to the reaper queue, such as the merge cleanup and decay services, backpressure additional requests when the queue is full. Previously, the reaper queue dropped events under heavy load. Added several new environment variables for tuning cleanup parallelism. One adds a manual release for catalogs with substantial inactive partition backlogs which allows the periodic service to clean up catalog data and dispatch reaper events. Expect changes in metrics related to object store request volume and partition counts.

Changes⚓︎

  • Partition reaping uses far fewer object store requests, now using bulk deletion. The reaper backs off if rate-limited. Failed deletes now drop the catalog record, leaving the files to the partition cleaner. New reaper_* metrics track throttled, deferred, and orphaned partitions.

  • The reaper now deletes a partition's known files directly instead of listing the partition prefix first, removing a LIST request per reaped partition. A reap that finds none of its expected files now warns and increments reaper_reap_all_not_found_total so orphaned keys (for example, after a storage bucket_path change) are visible. Set REAPER_LIST_BEFORE_DELETE to restore the previous list-then-delete behavior.

  • Improved reaper queue partition deletion efficiency. All tasks rejoin the back of the queue for an available worker preventing slow storage locations from blocking faster sibling storage locations. Faster storage locations claim more turns and delete faster.

  • Clusters with large partition backlogs clean up faster. The decay and merge-cleanup services now sweep tables in shuffled, interleaved batches, and merge-cleanup fetches locked partitions without sorting.

  • Adapted the decay and merge cleanup operations to test and respect RabbitMQ queue depth before enqueuing additional reaper events. By using back pressure to pace the producers, the system avoids dropping reap events on queue overrun. Also, the periodic service now emits metrics every minute instead of only once at the end of each batch run. This backlog observability improvement is necessary for bulk cleaning operations.

  • Improved throughput for collecting catalog entries to reap. Removed a catalog query bottleneck by using PostgreSQL SKIP LOCK and allowing multiple routines to concurrently collect disjoint rows for reaping from the catalog. Use environment variables DECAY_TABLE_CONCURRENCY and MERGE_CLEANUP_TABLE_CONCURRENCY to set maximum query parallelization.

  • Improved throughput for collecting catalog entries to reap. Added multiple in-process communication channels when sending reaper events to RabbitMQ. This removed a shared lock contention bottleneck, allowing parallel collectors to send reap events. It depends on the back pressure detection to avoid RabbitMQ overrun and reaper event drop.

  • Improved reaper system efficiency by dequeueing and acknowledging batches of reap events from RabbitMQ instead handling events one at a time. This increases queue throughput, decreases RabbitMQ contention, and keeps parallel reapers busy cleaning inactive partitions from storage.

  • Added a "delete on dispatch" mode to the decay and merge cleanup jobs in the periodic service. This mode helps drain a backlog of inactive partitions that the decay and merge cleanup jobs can't keep up with. When enabled, the query that finds inactive partitions also deletes their catalog rows and returns the necessary information for reaper to clean up the corresponding partition data. As a result, merge cleanup skips an additional "unlock update" step and the reaper no longer deletes the catalog rows later. If a reaper event is lost after its catalog row is deleted, the storage files are orphaned and partition cleaner reclaims them. Set DECAY_DELETE_ON_DISPATCH or MERGE_CLEANUP_DELETE_ON_DISPATCH to enable this mode. Both default to false.

Fixes⚓︎

  • Correctly propagate the billing bytes statistic into catalog metadata during merge. Before the fix, the merge system stored zero for billing_bytes for all merged partitions, though the originally ingested partitions had the correct value. The bug inhibited reporting and analysis, but didn't affected bills, as merged partitions aren't used for billing.

Observability⚓︎

Worth a look if you run dashboards or alerts on Hydrolix internals. Where spans come from changes this release, and a new metric family arrives. Panels pointing at the retired collectors go blank after the upgrade.

Changes⚓︎

  • Administrators can now see whether the hdx-pg-monitor service's PostgreSQL monitoring queries are succeeding and how long they take. New pgmonitor_* metrics report query executions, failures by cause, duration, and other diagnostic information.

  • Retired cluster-internal sidecar OpenTelemetry collectors in favor of a single central collector. The benefits of batched span collection aren't worth the increased cluster-wide resource cost.

  • Switched the Tulugaq metrics to transform management tools to use the v2 Config API endpoints.

Configuration and APIs⚓︎

The v2 Config API is feature-complete as of this release, so anything still on v1 can migrate. The v1 endpoints remain unchanged. Three Hydrolix cluster spec keys (s3_endpoint, premium_storage_class, and region) have been deprecated, and a fix ensures functions and UDFs aren't orphaned when a project is deleted.

Changes⚓︎

  • Migrated the last remaining Config API endpoints to v2, completing the v2 endpoint set.

  • Added parallel v2 Config API variants for most remaining non-nested v1 endpoints. Calling conventions don't change for the non-nested endpoints.

  • Added v2 Config API endpoints for column policies and row policies, managed as standalone resources rather than as part of the table definition.

  • Added v2 Config API endpoints for column and customer objects.

  • Added v2 Config API endpoints for the catalog and partitions at shorter, top-level paths that take the organization from the request rather than a path parameter. The equivalent v1 org-scoped endpoints are unchanged.

  • The Config API /customers endpoints can return a revision, a version number for the customer's configuration that changes whenever the config does, so clients can detect a change and skip reloading unchanged config. It's opt-in via ?include_revision=True because computing the revision adds overhead.

  • The Config API user response now includes a customers list, so a user's accessible customers can be read in a single call.

  • The s3_endpoint, premium_storage_class, and region spec keys are now deprecated tunables. Use db_bucket_endpoint in place of s3_endpoint, which remains a legacy fallback.

  • Improve the performance of a task that ensures users in a cluster who have a global permission are associated with all customer objects.

Fixes⚓︎

  • Deleting a project now also removes its functions and UDFs, which were previously left orphaned.

  • Fixed token issuance on the v2 service account endpoint, which was rejecting requests that didn't also supply name and org.

  • Corrected several hydrolix_otel tunable variables. Fixed incorrect descriptions of hdx_table and hdx_transform and updated the grpc_port and http_port defaults to 5317 and 5318, respectively.

Interfaces and tools⚓︎

Changes⚓︎

  • Upgraded the operator-bundled Hydrolix Grafana datasource plugin version to 0.11.0. Clusters that haven't pinned the version field use this plugin version up on upgrade. See the plugin's change log for full per-version change logs.

  • Updated the Kibana Gateway default version from v2.1.0 → v2.2.0. This defaults to the v2 Config API endpoints and supports the config_api_version of either v1 or v2 in the kibana_gateway_config tunable.

Fixes⚓︎

  • Fixed incorrect scale-value validation in the SIEM source creation and edit form.

Bundled component versions⚓︎

Third-party components whose bundled or default version changed in this release. Where a version change also changes behavior, the entry in the relevant area covers the detail.

Component Version in v6.4.x
Hydrolix Grafana datasource plugin 0.11.0
HTTP proxy (chproxy) v0.6.8
Kibana Gateway v2.2.0