This tutorial loads sample web access logs into Hydrolix and queries them. It takes about 15 minutes and uses only curl, jq, and the Hydrolix HTTP APIs.
Check that echo "${HDX_TOKEN}" prints a long string. If it prints null, the login failed. Run the command again without the last line to see the error.
Save the uuid of your organization and of the customer that should own the new project. If more than one customer is listed, ask your Hydrolix administrator which one to use.
Check that echo "${HDX_PROJECT}" prints a UUID. If it prints null, run the command again without the last line to see the error. If the error says getting_started already exists, a project from an earlier run is still there. Delete it as described in Clean up, then create the project again.
The sample data has fixed dates. By default, a table rejects rows more than 365 days old, and cold_data_max_age_days raises that limit so the sample loads whenever you run this tutorial. Tables that receive current data don't need this setting.
A transform tells Hydrolix how to read incoming data and which column holds the timestamp. Each event in the sample looks like this, with the user_agent value shortened.
Wait about a minute after creating the transform for the ingest service to load the new configuration. Then send the file to the streaming ingest endpoint.
A successful response is {"code":200,"message":"success"}. If the response is a 400 error saying a project or table wasn't found, or that the transform is unknown, wait another minute and send the file again. Hydrolix rejects the request before reading any events. Sending it again doesn't create duplicate rows.
Each query here runs through the query endpoint and filters on the timestamp, which lets Hydrolix skip data outside that day. New data is ready to query about a minute after it's sent. Until then, a query can fail with an error saying the table doesn't exist, or return \N and 0 in place of results. Wait a minute and run it again.
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT status, count() AS requests FROM getting_started.web_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' GROUP BY status ORDER BY requests DESC, status"
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT path, count() AS errors FROM getting_started.web_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' AND status >= 500 GROUP BY path ORDER BY errors DESC"
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT path, round(quantile(0.95)(response_time_ms), 1) AS p95_ms FROM getting_started.web_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' GROUP BY path ORDER BY p95_ms DESC LIMIT 5"
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT toStartOfInterval(timestamp, INTERVAL 4 HOUR) AS period, count() AS requests, countIf(status >= 500) AS errors FROM getting_started.web_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' GROUP BY period ORDER BY period"
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT status, count() AS requests FROM getting_started.web_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' GROUP BY status ORDER BY requests DESC, status FORMAT JSONEachRow"
Every query in this tutorial filters on timestamp. Without that filter, Hydrolix reads all of the table's data. With 200 rows that's fast, but on a large table it can take much longer.
Some clusters block queries that don't filter on time, using the hdx_query_timerange_required query option. To see what that looks like, turn it on for one query that has no filter.
hdx_query_timerange_required is set to true. Your query needs a time range filter in a WHERE clause.
Remove SETTINGS hdx_query_timerange_required = 1 and the jq filter, and the same query returns a count of 200.
A second sample file holds 200 requests from a content delivery network (CDN) on the same day. It has more fields than the web logs, including the edge location that served each request and whether the request was a cache hit. Load it into a second table in the same project.
curl-s-G"https://${HDX_HOSTNAME}/query"\-H"Authorization: Bearer ${HDX_TOKEN}"\--data-urlencode"query= SELECT cache_status, count() AS requests, round(avg(ttfb_ms), 1) AS avg_ttfb_ms FROM getting_started.cdn_logs WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00' GROUP BY cache_status ORDER BY requests DESC"
HIT 88 12.3
MISS 85 82
EXPIRED 19 88.9
STALE 8 73.3
In this sample, the average TTFB of a cache hit is about one-seventh of a miss. To see which edge locations miss the cache most often, group by edge_location and count misses with countIf(cache_status = 'MISS').