Pre-1.0 hoardDB is pre-1.0. Expect breaking changes.

Is this node full?

A node runs out of one of four things: file descriptors, memory mappings, memory, or disk. hoardDB measures all four against the limits your kernel and your filesystem actually enforce, and tells you which one is close. Read which resource is red before you act: the right action differs by resource, and for disk the usual reflex, adding a node, is the wrong one.

The one-sentence version: look at which resource is red. For fds or mappings, raise the kernel limit. For memory under load, add a node. For disk, free space on this node: a red node votes against new members, but in a cluster of three or more the others can still admit one, and a new node barely helps a full disk either way.

Where to look

You do not need to set anything up to be told. Three channels carry the same readings:

  • The node’s log. A WARN line is written when a resource crosses its line, and repeated at most once an hour while it stays there. Each line names the action for that resource. This is always on.

  • hoardDB status. Any authenticated session shows one row per resource:

    [h]oardDB> status
    STATUS
    ---
    ...
    Healthy:      true
    Capacity:
      fds         0.00  threshold 0.80  ok
      mappings    0.00  threshold 0.80  ok
      memory      0.00  threshold 0.80  ok
      disk        0.32  threshold 0.85  ok
    

    This is always on too.

  • Prometheus, at /metrics. Off by default; see Turning on the gauges.

What the states mean

statewhat it means
okbelow the threshold
approachingdisk only, with the high-water mark on: within 0.10 of it. Plan now
redat or over the threshold. Act on the row for that resource, below
unavailablenot fine. The resource could not be measured, so there is no ratio to show. The log says why, and what to check
stalenot fine. The last reading is more than 45 seconds old: whatever measures the node has stopped. The ratios shown are the last ones known, not current ones

Neither unavailable nor stale means healthy. A node that cannot be measured is a node you know nothing about.

A red row in status is followed by one line pointing you to the log, because the log line carries the action and, for memory, which of two things red means.

What to watch, with Prometheus

hoarddb_node_capacity_ratio{resource} is each resource’s reading divided by its live limit. hoarddb_node_capacity_threshold_ratio{resource} is where that resource turns red. The resource label takes exactly the values fds, mappings, memory and disk. Graph the four ratios with:

max by (resource) (hoarddb_node_capacity_ratio)

Alert on three rules, not one. The first alone is silent when a resource stops being measured, because the ratio then disappears rather than reading zero:

# a resource is over its threshold
hoarddb_node_capacity_ratio > on(resource) hoarddb_node_capacity_threshold_ratio
# a resource can no longer be measured (the threshold is always published; the ratio is not)
hoarddb_node_capacity_threshold_ratio unless on(resource) hoarddb_node_capacity_ratio
# the sampler has stopped
time() - hoarddb_node_capacity_last_sample_timestamp_seconds > 60

The thresholds

The thresholds are policy lines, not measured cliffs. They are where this release chooses to warn you, not points at which hoardDB has been measured to fail. Where your node actually degrades depends on your hardware and workload, and no figure for it is published yet.

resourcered atamber
fds0.80none
mappings0.80none
memory0.80none
diskdisk_high_water_pct ÷ 100 (default 0.85), or 0.80 when that is 00.10 below red, only while disk_high_water_pct is on

None of these is configurable except through disk_high_water_pct, which already exists for another reason.

Disk and disk_high_water_pct

server.disk_high_water_pct (HOARDB_DISK_HIGH_WATER, default 85) is the point at which this node votes against new members. The disk warning fires at the same point, from the same measurement, so the warning and the vote always agree.

  • 0 turns off only the refusal. The node no longer votes against joins or refuses to propose one, but the disk still warns, at 0.80: turning off the vote does not turn off being told your disk is filling. There is no amber line in this mode, because there is no vote to warn you ahead of.
  • A value above 100 turns off both the refusal and the warning. The threshold is then above anything a disk can reach.

When a resource is red

Read the disk row first. It is the one where the obvious action is wrong.

resource redwhat it usually meansdo thisdoes adding a node help?
diskThe data directory’s filesystem is filling up: the store’s own bytes, a write-ahead log held back for a replica that is unreachable, blob payloads, or something that is not hoardDB.Free space on this node. (1) Compare hoarddb_node_data_bytes with hoarddb_node_disk_used_bytes to see whether hoardDB is the consumer at all. (2) If hoarddb_node_data_bytes{part="wal"} is large, bring the unreachable peer back or remove it from the cluster: the log is kept for it on purpose. (3) If hoarddb_reclaim_pending_records is above 0, dropped data is still being reclaimed, and its space returns only after compaction. (4) Grow the volume, or drop buckets or databases you no longer need. Do not grow the cluster while this node is red.No. See below.
disk, approachingas above, earlierPlan now. If growth is organic and even across buckets, add capacity while every node is below the mark, so every node votes for the join. If one bucket or the log is growing, treat it as the red row.Yes, as planning, while every node’s disk is below the mark
fdsDescriptors are held mostly by the store’s table and value-log files, which grow with stored bytes, plus client and peer connections (server.max_conns and server.max_peer_conns)Raise the process’s RLIMIT_NOFILE: LimitNOFILE= in a systemd unit, --ulimit nofile= in Docker. It is a kernel limit, and raising it is free. If it is already very high and still close, the node holds a lot of data: read the disk row.Not as the first action
mappingsThe store maps its table files into memory, so mappings also grow with stored bytesRaise vm.max_map_count: sysctl -w vm.max_map_count=262144, and persist it in /etc/sysctl.d/. The kernel default of 65530 is low for any database that maps its files.Not as the first action
memoryAnonymous memory, resident or swapped: the block and index caches (storage.badger_block_cache_mb and badger_index_cache_mb), per-request working memory, the in-memory bucket registry, and a rebalance in progress, which holds what it moves in memory on both nodes. Memory is red: read the log’s memory-limit line first; it tells you which of two things red means. Against memory.max or the host’s RAM, the node is near where it swaps heavily or is killed, so act now. Against memory.high, the kernel is already throttling this process. The node is slow, not about to be killed. memory.high is a throttle line, not a hard limit(1) If a rebalance is running, wait for it; the peak passes. (2) If the caches were raised, lower them. (3) If the load is traffic-driven, add a node, and expect a brief rise on this node while buckets move away. (4) Against memory.max or host RAM: if the limit is below the host’s memory, raise it. Against memory.high: raise it if it was set lower than you meant (MemoryHigh= in systemd), and add a node if the load is real.Yes, when traffic or the number of buckets drives it. It does not shrink the fixed caches, and the rebalance itself costs memory first

Why adding a node does not fix a full disk

Disk is red: free space on this node. Adding a node is not the fix. At the high-water mark this node votes against new members. In a one- or two-node cluster that blocks the join. In a cluster of three or more, the other nodes can still admit one, and this node then takes part in the rebalance, so do not grow the cluster while a node’s disk is red. Either way, a new node frees little space on this one, and slowly: buckets move by hash and not by size, freed space comes back only after cleanup and compaction, and space held by values of 1 MiB or more is not returned at all in this release. Add nodes while every node’s disk is green, as planning.

The log line for a red disk says the same:

disk on this node is at 0.91 of its filesystem (high-water mark 0.85): free space on this node or grow the volume. At this mark the node votes against joins, but in a cluster of three or more a majority can still admit one, and this node then takes part in the rebalance. Do not grow the cluster while this node is red

and, with disk_high_water_pct on, the early warning:

disk on this node is at 0.78, approaching the high-water mark 0.85: free space or grow the volume now. At the mark this node votes against joins

What the memory ratio counts

The memory ratio is anonymous memory, resident plus swapped (RssAnon + VmSwap from /proc/<pid>/status), over the memory limit. Page cache does not count: the kernel sheds it under pressure.

Swapped memory counts as used, so a node that is swapping reads high and can read above 1.0. Against memory.max or host RAM, above 1.0 means the node is over its hard limit and swapping. Against memory.high, it means the node is past the line where the kernel throttles it.

The limit is the smallest memory.max or memory.high on the process’s cgroup and every parent cgroup it can see, or the host’s MemTotal when none is set. The log says which one it is using, when the node starts and again whenever it changes:

memory limit 4.0 GiB from memory.max on cgroup /system.slice
memory limit 2.0 GiB from memory.high on cgroup /system.slice/hoarddb.service: this is a throttle line, not a hard limit. Past it the kernel slows this process and reclaims its memory; it does not kill it there. The hard limit is 4.0 GiB from memory.max on cgroup /system.slice

The two red lines read differently, on purpose:

memory on this node is at 0.86 of its memory limit 4.0 GiB (memory.max on cgroup /system.slice): anonymous memory, resident plus swapped, is near the point where this process swaps heavily or is killed. If a rebalance is running, wait for it. If the block or index caches were raised, lower them. If the load is traffic-driven, add a node. If the limit is below the host's memory, raise it.
memory on this node is at 0.86 of memory.high 2.0 GiB (cgroup /system.slice/hoarddb.service), the line where the kernel starts throttling this process and reclaiming its memory. The node is being slowed, not stopped: memory.high is not a hard limit, and the process is not killed there. If memory.high was set lower than you meant (MemoryHigh= in systemd), raise it. If a rebalance is running, wait for it. If the block or index caches were raised, lower them. If the load is traffic-driven, add a node.

Inside a container, set the memory limit on the container itself. A process in its own cgroup namespace, which is Docker’s default, cannot see a limit set above the container, for example only on a pod. The log then says the limit in use is the host’s MemTotal and that a limit above the container is not seen.

On a host that manages memory with cgroup v1, hoardDB does not read the v1 limit. It uses the host’s MemTotal and writes a warning at start and every hour saying so; if the process runs under a v1 memory limit, the memory ratio reads lower than it should.

Turning on the gauges

The gauges are exported only when /metrics is on. It is a deploy-time setting:

HOARDB_METRICS_ENABLED=true HOARDB_METRICS_LISTEN=10.0.0.5:9090 hoardDB-server

The log warnings and status work without it. When it is off, the node says so once at start:

capacity gauges are not exported (HOARDB_METRICS_ENABLED is false); capacity warnings still go to this log and to hoardDB status

/metrics is unauthenticated. It listens on 0.0.0.0:9090 by default. Bind it to a private interface, as above, or firewall the port.

The series, all node-wide (none is per bucket, per database or per peer):

serieswhat it is
hoarddb_node_capacity_ratio{resource}reading ÷ limit. Absent while the resource cannot be measured
hoarddb_node_capacity_threshold_ratio{resource}where that resource turns red. Always present
hoarddb_node_capacity_sample_errors_total{resource}samples in which the resource could not be read
hoarddb_node_capacity_last_sample_timestamp_secondswhen the last sample completed
hoarddb_node_bucketsbuckets this node serves
hoarddb_node_memory_mappings, hoarddb_node_memory_mappings_limitmappings, and vm.max_map_count
hoarddb_node_memory_anon_bytes, hoarddb_node_memory_limit_bytesthe memory ratio’s two sides
hoarddb_node_disk_used_bytes, hoarddb_node_disk_size_bytesthe disk ratio’s two sides, from one measurement
hoarddb_node_data_bytes{part}what hoardDB itself holds: store_lsm, store_vlog (up to a minute old) and wal. wal is absent on a single node, which keeps no log
process_open_fds, process_max_fdsthe descriptor ratio’s two sides

Reading your node’s limits

Every ratio divides by a limit you can read yourself. Replace <pid> with the server’s process id (pgrep hoardDB-server) and <data_dir> with its data directory.

# descriptors: the "Max open files" soft limit is the fds limit
grep 'Max open files' /proc/<pid>/limits
ls /proc/<pid>/fd | wc -l

# mappings
sysctl vm.max_map_count
wc -l < /proc/<pid>/maps

# disk
df -h <data_dir>

# memory: the limits on the process's cgroup and every parent it can see
cg=$(sed -n 's/^0:://p' /proc/<pid>/cgroup)
while [ -n "$cg" ]; do echo "$cg  memory.max=$(cat /sys/fs/cgroup$cg/memory.max 2>/dev/null)  memory.high=$(cat /sys/fs/cgroup$cg/memory.high 2>/dev/null)"; cg=${cg%/*}; done
grep -E 'RssAnon|VmSwap' /proc/<pid>/status

max means no limit at that level. The walk assumes cgroup v2 mounted at /sys/fs/cgroup, which is the common layout; hoardDB itself finds the mount through /proc/<pid>/mountinfo.

Not in the capacity set

  • CPU is a watch signal, not a ratio: rate(process_cpu_seconds_total[5m]) against the node’s core count. Sustained, request-driven CPU is the classic case where adding a node helps. CPU during a rebalance, a reclaim or a compaction is expected, and passes.
  • process_network_receive_bytes_total and process_network_transmit_bytes_total do not measure hoardDB. They read the network namespace’s counters, so on a host with shared networking they report the whole machine’s traffic. Use hoarddb_request_duration_seconds to see whether the network is slowing requests down.

Source: docs/user/capacity.md in the repository.