Skip to content

Get Started

This tutorial loads sample web access logs into Hydrolix and queries them. It takes about 15 minutes and uses only curl, jq, and the Hydrolix HTTP APIs.

What you'll do⚓︎

Each step uses values you export in the steps before it.

  1. Get an API token.
  2. Find your organization and customer.
  3. Create a project, a table, and a transform.
  4. Send the sample data.
  5. Query it.

Before you begin⚓︎

You need:

  • The hostname of a Hydrolix cluster, such as hostname.hydrolix.live, and an account on it that can create projects
  • curl
  • jq, which the commands use to pull values out of responses
  • A Bash-compatible shell, such as the macOS or Linux terminal, or Windows Subsystem for Linux (WSL)

Set the hostname once, so the commands on this page can use it.

Set the Cluster Hostname
export HDX_HOSTNAME=hostname.hydrolix.live

Get an API token⚓︎

Set your username, which is your email address. Then enter your password at the prompt, which keeps it out of your shell history.

Set Your Credentials
export HDX_USER=<email address>
read -s HDX_PASSWORD && export HDX_PASSWORD

Log in to get a bearer token. By default, the token is valid for 24 hours.

Log In to the API
1
2
3
4
export HDX_TOKEN=$(curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/login/" \
  -H "Content-Type: application/json" \
  -d "$(jq -n --arg u "${HDX_USER}" --arg p "${HDX_PASSWORD}" '{username: $u, password: $p}')" \
  | jq -r '.auth_token.access_token')

Check that echo "${HDX_TOKEN}" prints a long string. If it prints null, the login failed. Run the command again without the last line to see the error.

For other ways to get a token, including service accounts, see Authenticate to the API.

Find your organization and customer⚓︎

Hydrolix groups projects under an organization and a customer. List the ones your account belongs to.

List Your Organization and Customers
1
2
3
curl -s "https://${HDX_HOSTNAME}/config/v2/users/current/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  | jq '{orgs: [.orgs[] | {uuid, name}], customers: [.customers[] | {uuid, name}]}'

Save the uuid of your organization and of the customer that should own the new project. If more than one customer is listed, ask your Hydrolix administrator which one to use.

Save the Organization and Customer
export HDX_ORG=<org uuid>
export HDX_CUSTOMER=<customer uuid>

Create a project⚓︎

A project groups related tables. Create one named getting_started.

Create a Project
1
2
3
4
5
export HDX_PROJECT=$(curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/projects/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{"name": "getting_started", "org": "'"${HDX_ORG}"'", "customer": "'"${HDX_CUSTOMER}"'"}' \
  | jq -r '.uuid')

Check that echo "${HDX_PROJECT}" prints a UUID. If it prints null, run the command again without the last line to see the error. If the error says getting_started already exists, a project from an earlier run is still there. Delete it as described in Clean up, then create the project again.

Create a table⚓︎

A table holds the data. Create one named web_logs in the new project.

Create a Table
1
2
3
4
5
6
7
8
9
export HDX_TABLE=$(curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/tables/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "web_logs",
    "project": "'"${HDX_PROJECT}"'",
    "settings": {"stream": {"cold_data_max_age_days": 3650}}
  }' \
  | jq -r '.uuid')

The sample data has fixed dates. By default, a table rejects rows more than 365 days old, and cold_data_max_age_days raises that limit so the sample loads whenever you run this tutorial. Tables that receive current data don't need this setting.

Create a transform⚓︎

A transform tells Hydrolix how to read incoming data and which column holds the timestamp. Each event in the sample looks like this, with the user_agent value shortened.

One Sample Event
{"timestamp":"2026-09-01T00:13:39Z","client_ip":"203.0.113.41","method":"GET","path":"/api/search","status":200,"bytes_sent":6774,"response_time_ms":34.7,"user_agent":"Mozilla/5.0 ...","country":"US"}

Create a transform that maps each field to a column, with timestamp as the primary column.

Create a Transform
curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/transforms/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "web_logs_json",
    "type": "json",
    "table": "'"${HDX_TABLE}"'",
    "settings": {
      "is_default": true,
      "output_columns": [
        {"name": "timestamp", "datatype": {"type": "datetime", "format": "2006-01-02T15:04:05Z", "primary": true, "resolution": "seconds"}},
        {"name": "client_ip", "datatype": {"type": "string"}},
        {"name": "method", "datatype": {"type": "string"}},
        {"name": "path", "datatype": {"type": "string"}},
        {"name": "status", "datatype": {"type": "uint32"}},
        {"name": "bytes_sent", "datatype": {"type": "uint64"}},
        {"name": "response_time_ms", "datatype": {"type": "double"}},
        {"name": "user_agent", "datatype": {"type": "string"}},
        {"name": "country", "datatype": {"type": "string"}}
      ]
    }
  }' \
  | jq '{uuid, name}'

The format value is a Go time layout. It describes how the timestamps in the data are written. For other types and formats, see Basic Data Types.

Send the sample data⚓︎

Download the sample file. It holds 200 web requests from one day, one JSON event per line. All data in the sample is synthetic.

Download the Sample Data
curl -fsSO https://docs.hydrolix.io/latest/assets/files/sample-data/web-access-logs.json

Wait about a minute after creating the transform for the ingest service to load the new configuration. Then send the file to the streaming ingest endpoint.

Send the Sample Data
1
2
3
4
5
6
curl -s -X POST "https://${HDX_HOSTNAME}/ingest/event" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -H "X-Hdx-Table: getting_started.web_logs" \
  -H "X-Hdx-Transform: web_logs_json" \
  --data-binary @web-access-logs.json

A successful response is {"code":200,"message":"success"}. If the response is a 400 error saying a project or table wasn't found, or that the transform is unknown, wait another minute and send the file again. Hydrolix rejects the request before reading any events. Sending it again doesn't create duplicate rows.

Query your data⚓︎

Each query here runs through the query endpoint and filters on the timestamp, which lets Hydrolix skip data outside that day. New data is ready to query about a minute after it's sent. Until then, a query can fail with an error saying the table doesn't exist, or return \N and 0 in place of results. Wait a minute and run it again.

Count the requests by status code.

Requests by Status Code
1
2
3
4
5
6
7
8
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT status, count() AS requests
    FROM getting_started.web_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
    GROUP BY status
    ORDER BY requests DESC, status"
Requests by Status Code: Result
1
2
3
4
200  179
404  13
304  4
500  4

Find which paths returned server errors.

Server Errors by Path
1
2
3
4
5
6
7
8
9
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT path, count() AS errors
    FROM getting_started.web_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
      AND status >= 500
    GROUP BY path
    ORDER BY errors DESC"
Server Errors by Path: Result
/checkout  4

Every server error came from /checkout. Check which paths are slowest, by 95th percentile response time.

Slowest Paths
1
2
3
4
5
6
7
8
9
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT path, round(quantile(0.95)(response_time_ms), 1) AS p95_ms
    FROM getting_started.web_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
    GROUP BY path
    ORDER BY p95_ms DESC
    LIMIT 5"
Slowest Paths: Result
1
2
3
4
5
/checkout     158.8
/             145.2
/products     141
/api/orders   134.1
/products/42  127.4

/checkout also has the highest 95th percentile response time.

More to try⚓︎

Learn more about queries, transforms, and CDN logs.

Group requests into 4-hour periods to see how traffic and errors change over the day. For hourly periods, use toStartOfHour(timestamp) instead.

Requests per Period
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT toStartOfInterval(timestamp, INTERVAL 4 HOUR) AS period,
           count() AS requests,
           countIf(status >= 500) AS errors
    FROM getting_started.web_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
    GROUP BY period
    ORDER BY period"
Requests per Period: Result
1
2
3
4
5
6
2026-09-01 00:00:00  20  1
2026-09-01 04:00:00  44  1
2026-09-01 08:00:00  29  0
2026-09-01 12:00:00  30  1
2026-09-01 16:00:00  44  0
2026-09-01 20:00:00  33  1

Results come back as tab-separated text by default. Add FORMAT JSONEachRow to get one JSON object per row, which is easier for a script to read.

Requests by Status Code as JSON
1
2
3
4
5
6
7
8
9
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT status, count() AS requests
    FROM getting_started.web_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
    GROUP BY status
    ORDER BY requests DESC, status
    FORMAT JSONEachRow"
Requests by Status Code as JSON: Result
1
2
3
4
{"status":200,"requests":179}
{"status":404,"requests":13}
{"status":304,"requests":4}
{"status":500,"requests":4}

Every query in this tutorial filters on timestamp. Without that filter, Hydrolix reads all of the table's data. With 200 rows that's fast, but on a large table it can take much longer.

Some clusters block queries that don't filter on time, using the hdx_query_timerange_required query option. To see what that looks like, turn it on for one query that has no filter.

Query Without a Time Filter
1
2
3
4
5
6
7
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT count()
    FROM getting_started.web_logs
    SETTINGS hdx_query_timerange_required = 1" \
  | jq -r '.error'

The error includes this message, along with more detail about which filters count.

Query Without a Time Filter: Result (Shortened)
hdx_query_timerange_required is set to true. Your query needs a time range filter in a WHERE clause.

Remove SETTINGS hdx_query_timerange_required = 1 and the jq filter, and the same query returns a count of 200.

A second sample file holds 200 requests from a content delivery network (CDN) on the same day. It has more fields than the web logs, including the edge location that served each request and whether the request was a cache hit. Load it into a second table in the same project.

Download the CDN Sample Data
curl -fsSO https://docs.hydrolix.io/latest/assets/files/sample-data/cdn-logs.json
Create the CDN Table
1
2
3
4
5
6
7
8
9
export HDX_CDN_TABLE=$(curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/tables/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "cdn_logs",
    "project": "'"${HDX_PROJECT}"'",
    "settings": {"stream": {"cold_data_max_age_days": 3650}}
  }' \
  | jq -r '.uuid')
Create the CDN Transform
curl -s -X POST "https://${HDX_HOSTNAME}/config/v2/transforms/" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "cdn_logs_json",
    "type": "json",
    "table": "'"${HDX_CDN_TABLE}"'",
    "settings": {
      "is_default": true,
      "output_columns": [
        {"name": "timestamp", "datatype": {"type": "datetime", "format": "2006-01-02T15:04:05Z", "primary": true, "resolution": "seconds"}},
        {"name": "edge_location", "datatype": {"type": "string"}},
        {"name": "client_ip", "datatype": {"type": "string"}},
        {"name": "host", "datatype": {"type": "string"}},
        {"name": "method", "datatype": {"type": "string"}},
        {"name": "path", "datatype": {"type": "string"}},
        {"name": "protocol", "datatype": {"type": "string"}},
        {"name": "status", "datatype": {"type": "uint32"}},
        {"name": "cache_status", "datatype": {"type": "string"}},
        {"name": "bytes_sent", "datatype": {"type": "uint64"}},
        {"name": "ttfb_ms", "datatype": {"type": "double"}},
        {"name": "origin_time_ms", "datatype": {"type": "double"}},
        {"name": "country", "datatype": {"type": "string"}}
      ]
    }
  }' \
  | jq '{uuid, name}'

Wait about a minute, then send the file.

Send the CDN Sample Data
1
2
3
4
5
6
curl -s -X POST "https://${HDX_HOSTNAME}/ingest/event" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  -H "Content-Type: application/json" \
  -H "X-Hdx-Table: getting_started.cdn_logs" \
  -H "X-Hdx-Transform: cdn_logs_json" \
  --data-binary @cdn-logs.json

Compare time to first byte (TTFB) for cache hits and misses.

TTFB by Cache Status
1
2
3
4
5
6
7
8
curl -s -G "https://${HDX_HOSTNAME}/query" \
  -H "Authorization: Bearer ${HDX_TOKEN}" \
  --data-urlencode "query=
    SELECT cache_status, count() AS requests, round(avg(ttfb_ms), 1) AS avg_ttfb_ms
    FROM getting_started.cdn_logs
    WHERE timestamp >= '2026-09-01 00:00:00' AND timestamp < '2026-09-02 00:00:00'
    GROUP BY cache_status
    ORDER BY requests DESC"
TTFB by Cache Status: Result
1
2
3
4
HIT      88  12.3
MISS     85  82
EXPIRED  19  88.9
STALE     8  73.3

In this sample, the average TTFB of a cache hit is about one-seventh of a miss. To see which edge locations miss the cache most often, group by edge_location and count misses with countIf(cache_status = 'MISS').

Clean up⚓︎

To remove what this tutorial created, delete the project. Deleting it also deletes its tables, transforms, and data.

Delete the Project
curl -s -X DELETE "https://${HDX_HOSTNAME}/config/v2/projects/${HDX_PROJECT}/" \
  -H "Authorization: Bearer ${HDX_TOKEN}"

Next steps⚓︎