Francis
GitHub

Jobs

A job is a durable, fire-and-forget task dispatched to a specific actor. A client (an external app, or another actor) dispatches a job to an actor (type, id), which Francis then runs in the background on whatever host serves that actor type, delivered to the actor’s Job method. Jobs are stored in the database, so they survive restarts, are retried automatically, and on permanent failure are dead-lettered (recorded for inspection and replay) rather than lost.

Jobs run immediately, after a delay, on a repeating interval, or on a cron schedule. They take input and return nothing.

Jobs vs. alarms#

Jobs and alarms both ride on the same durable scheduling engine, but they serve different purposes:

AlarmsJobs
Delivered toAlarm(name, data)Job(method, data)
Keyed byname (set replaces on collision)server-issued JobID (each dispatch is distinct)
Dispatching NN alarms with the same name coalesce to oneN dispatches run N times
On permanent failurethe alarm is deletedthe job is dead-lettered
On successthe alarm is deletedthe job is either deleted or retained for auditing, if the actor type asks for it

Use an alarm for a self-scheduled reminder or timer that an actor owns (one per name). Use a job to dispatch background work to an actor, especially when you need each dispatch to run and failures to be recorded rather than dropped.

To distribute many independent tasks across a cluster — running a bounded number per host and scaling out with more hosts — use the built-in task pool , which builds this pattern on top of jobs.

Receiving jobs#

To receive jobs, an actor must implement the actor.ActorJob interface:

func (w *Worker) Job(ctx context.Context, method string, data actor.Envelope) error {
	// "method" identifies which job to run
	// "data" is the optional payload attached when the job was dispatched (may be nil)
	switch method {
	case "send-email":
		// ... send the email ...
	}
	return nil
}
  • method is the value passed at dispatch time. An actor can handle many job methods.
  • data is an actor.Envelope. Call data.Decode(&dest) to read the payload. It may be nil.
  • Returning an error retries the job per the actor type’s MaxAttempts / InitialRetryDelay. Once retries are exhausted the job is dead-lettered.
  • Returning actor.ErrJobPermanentFailure skips the remaining retries and dead-letters the job immediately.

Dispatching a job#

Use the client (from inside an actor) to dispatch to the current actor:

payload := map[string]any{"to": "user@example.com"}
jobID, created, err := w.client.Dispatch(ctx, "send-email", payload)

From outside an actor, use the service, which targets any actor:

jobID, created, err := service.Dispatch(ctx, "worker", "worker-7", "send-email", payload)

Dispatch returns a server-issued jobID that is globally unique. Each call dispatches a distinct job: dispatching the same method twice runs it twice, unless you supply an idempotency key (in which case only the call that created the job gets created as true) - see below.

Job options#

Scheduling is controlled with actor.JobOption values:

OptionDescription
WithJobDueTime(t)Run first at an absolute time. The default is as soon as possible.
WithJobDelay(d)Run first after a delay, relative to dispatch. Mutually exclusive with WithJobDueTime.
WithJobInterval(iso8601)Repeat on an ISO 8601 / Go-duration interval. Mutually exclusive with WithJobCron.
WithJobCron(expr)Repeat on a standard cron expression. Mutually exclusive with WithJobInterval.
WithJobTTL(t)Stop a repeating job after this time.
WithIdempotencyKey(key)Dedup re-dispatch within the same actor (first dispatch wins).
// Run in 5 minutes
jobID, _, err := w.client.Dispatch(ctx, "reconcile", payload, actor.WithJobDelay(5*time.Minute))

// Run every weekday at 9am
jobID, _, err := w.client.Dispatch(ctx, "daily-report", nil, actor.WithJobCron("0 9 * * 1-5"))

Idempotency keys#

A job ID is server-issued and always unique. An idempotency key is caller-supplied and scoped to the actor: dispatching twice with the same key produces a single job (the first wins), so a retry of the dispatch itself does not enqueue duplicate work.

// Re-dispatching with the same key returns the same job, and only the first call reports created
// The work runs once
jobID, created, err := w.client.Dispatch(ctx, "charge", payload, actor.WithIdempotencyKey("order-42"))

Without a key, every dispatch is a distinct job: this is the anti-coalescing guarantee that distinguishes jobs from alarms.

When a job ends#

A job that ends (whether it completed or failed terminally) can leave a record behind. Both kinds of record live in the same store and are read back with the same calls.

When a job exhausts its retries, or returns actor.ErrJobPermanentFailure, it is dead-lettered. You can inspect, replay, or delete dead jobs (see below).

The two kinds are retained independently, because they are worth keeping for different reasons and for different lengths of time:

host.RegisterActor("worker", newWorker,
	// Keep a record of the runs that succeeded for a day
	local.WithCompletedJobRetention(24*time.Hour),
	// And the record of one that failed for a quarter
	local.WithDeadLetteredJobRetention(90*24*time.Hour),
)
OptionDefaultWhat the default means
WithCompletedJobRetention0A completed job leaves nothing behind by default
WithDeadLetteredJobRetention30 daysA dead-lettered job is recorded and readable for 30 days

A dead-lettered job is always recorded — dropping a failure silently is never useful — so its option only decides for how long.
A completed job is recorded only when you ask, since only you know whether a successful run’s history is worth storing.
Either option takes a negative duration to mean “keep the record with no expiry at all”.

A completed record carries the job’s metadata but not its payload. Only a dead job is ever replayed, so only a dead job keeps the data a replay would need.

Built-in actors keep their successes by default.
Every built-in gets 24 hours unless it asks for its own window, while a cron job keeps 7 days.
Their dead-lettered records take the same 30-day default as anything else.
This is for built-in actors only: an application actor still defaults to keeping none of its successes.

Optionally, an actor can react to a dead-lettering by implementing actor.ActorJobFailed:

func (w *Worker) JobFailed(ctx context.Context, jobID string, method string, data actor.Envelope, jobErr error) error {
	// Best-effort hook called after the job has been recorded in the dead-letter store
	// The dead-letter record is the source of truth: an error returned here is only logged
	return nil
}

The hook is best-effort: the dead-letter record is written first, then the hook is delivered on the actor’s turn-based lock. An error from the hook is logged and does not re-trigger dead-lettering.

For a repeating job, a single terminally-failing occurrence is dead-lettered while the recurrence continues: later occurrences still run.

Managing jobs#

GetJob, ListJobs, DeleteJob, and RetryJob are available on both the client (self-bound) and the service (for any actor):

// Inspect a job, spanning live and ended jobs alike
info, err := service.GetJob(ctx, jobID)
// info.Status is one of actor.JobStatusPending, JobStatusActive, JobStatusCompleted, JobStatusDeadLettered
// info.Status.IsTerminal() reports whether the job has ended

// List an actor's jobs: the live ones, plus any retained record
jobs, err := service.ListJobs(ctx, "worker", "worker-7")

// Remove a job, whatever state it is in
// One still scheduled is cancelled before it runs
// One that has ended has its record removed
err = service.DeleteJob(ctx, "worker", "worker-7", jobID)

// Re-dispatch a dead-lettered job (runs as soon as possible) and remove the dead record
newJobID, err := service.RetryJob(ctx, jobID)

GetJob returns actor.JobInfo. Attempts and EndedAt are populated only once a job has ended, and LastError only for one that dead-lettered. Missing jobs are reported as actor.ErrJobNotFound.

Delivery semantics#

  • Durable: jobs are persisted, so they survive restarts and are delivered after downtime.
  • Leased execution: hosts (or the runtime, in remote topologies) lease a job before running it, so it isn’t executed by two hosts at once.
  • Activation on dispatch: when a job is due, Francis activates the target actor if it isn’t already active, then calls Job.
  • Turn-based serialization: Job, Alarm, and Invoke all share the actor’s single turn lock, so a job never runs concurrently with other work on the same actor.
  • At-least-once with retries: a failing Job is retried per the actor type’s MaxAttempts and InitialRetryDelay, then dead-lettered. Design job handlers to be idempotent.

Full example#

func (w *Worker) Invoke(ctx context.Context, method string, data actor.Envelope) (any, error) {
	switch method {
	case "enqueue-report":
		// Dispatch background work to self
		// It runs immediately on whatever host serves this actor
		jobID, _, err := w.client.Dispatch(ctx, "report", map[string]any{"day": "today"})
		if err != nil {
			return nil, err
		}
		return map[string]string{"jobID": jobID}, nil
	}
	return nil, nil
}

func (w *Worker) Job(ctx context.Context, method string, data actor.Envelope) error {
	if method == "report" {
		// ... build and store the report ...
	}
	return nil
}

func (w *Worker) JobFailed(ctx context.Context, jobID string, method string, data actor.Envelope, jobErr error) error {
	// ... alert that a report job was dead-lettered ...
	return nil
}
Edit this page on GitHub