Nomad

A focused client for the HashiCorp Nomad HTTP API: inspect, run, scale and stop jobs.

Import
github.com/HiWay-Media/hwm-go-utils/nomad
API
/v1/… endpoints of any Nomad agent (default port 4646)
Auth
Optional ACL token (X-Nomad-Token)
  1. Setup
  2. Operations
  3. Examples
    1. Clone a job with a new name
    2. Scale out, then back
    3. List what runs on a node
    4. Read live resource usage

Setup

nc := nomad.NewService(nomad.Options{
	BaseUrl:    "http://nomad.service.consul:4646", // with or without a trailing /v1
	Token:      os.Getenv("NOMAD_TOKEN"),            // optional, for ACL-enabled clusters
	ScaleGroup: "encoder",                           // task group used by ScaleJob / RestartJob
	Logger:     logger,                              // optional, defaults to a no-op logger
	LogLevel:   "info",                              // "debug" logs every HTTP exchange
})
Option Default Notes
BaseUrl — http://host:4646 and http://host:4646/v1 are equivalent
Token none Sent as X-Nomad-Token on every request
ScaleGroup nomad.DefaultScaleGroup ("restreamer") Task group targeted by scaling
Logger no-op *zap.SugaredLogger
LogLevel — "debug" enables resty request/response logging

Operations

Method Nomad endpoint Returns
GetDefinition(jobID, region) GET /v1/job/:id *JobDefinition
RunJob(definition, region) POST /v1/jobs error
ScaleJob(jobID, count, region) POST /v1/job/:id/scale error
RestartJob(jobID, region) scale to 0, then back to 1 after one second error (of the first step)
DeleteJob(jobID, region, purge) DELETE /v1/job/:id?purge= error
AllocationStats(allocID, region) GET /v1/client/allocation/:id/stats *ResourceUsage
GetAllocations(nodeID, region) GET /v1/node/:id/allocations *NomadAllocations

Job, allocation and node ids are path-escaped, so ids containing spaces or slashes are safe.

Every call passes region as a query parameter; use "" for the agent’s default region.

Examples

Clone a job with a new name

def, err := nc.GetDefinition("restreamer-template", "eu")
if err != nil {
	return err
}
def.ID, def.Name = "restreamer-match-42", "restreamer-match-42"
return nc.RunJob(*def, "eu")

Scale out, then back

if err := nc.ScaleJob("restreamer-match-42", 3, "eu"); err != nil {
	return err
}
// …
return nc.ScaleJob("restreamer-match-42", 1, "eu")

List what runs on a node

allocs, err := nc.GetAllocations(nodeID, "eu")
if err != nil {
	return err
}
for _, a := range allocs.NomadAllocations {
	fmt.Printf("%s  %s/%s\n", a.ID, a.JobID, a.TaskGroup)
}

Read live resource usage

usage, err := nc.AllocationStats(allocID, "eu")
if err != nil {
	return err
}
fmt.Printf("RSS %d MiB, CPU %.1f%%\n", usage.MemoryStats.RSS>>20, usage.CpuStats.Percent)

RestartJob returns as soon as the job is scaled to 0; the scale back to 1 happens in the background one second later and its error is only logged. Call ScaleJob twice yourself if you need to wait for, or check, the second step.