Why Go's Concurrency Primitives Aren't Always Enough

Go's channels and goroutines are often held up as the gold standard for language-level concurrency. They strike a balance between power and flexibility, avoiding the deadlock-prone pitfalls of pthreads, the maintenance burden of callbacks, and the conceptual overhead of promises.

Yet there's a recurring gap in that toolkit. Time and again, Go developers find themselves hand-rolling a worker pool: a fixed number of m goroutines consuming n tasks from a channel-based queue. Goroutines are cheap enough that the classic rationale for thread pools—amortizing thread creation—doesn't apply. The real value is throttling concurrency when tasks saturate external resources like I/O or remote APIs.

The canonical implementation is deceptively simple:

package main

import (
	"fmt"
	"time"
)

func worker(id int, jobs <-chan int, results chan<- int) {
	for j := range jobs {
		fmt.Println("worker", id, "processing job", j)
		time.Sleep(time.Second)
		results <- j * 2
	}
}

func main() {
	jobs := make(chan int, 100)
	results := make(chan int, 100)

	for w := 1; w <= 3; w++ {
		go worker(w, jobs, results)
	}

	for j := 1; j <= 9; j++ {
		jobs <- j
	}
	close(jobs)

	for a := 1; a <= 9; a++ {
		<-results
	}
}

Three workers process nine jobs, each sleeping to simulate work. Closing the jobs channel after enqueueing everything signals the workers to exit their range loops. The elegance is the point: because this pattern is so trivial to assemble from primitives, there's no need for a standard library utility.

That's true—until real-world requirements arrive.

Where the Simple Model Breaks Down

The textbook example assumes every job succeeds. Production systems must handle errors. Adding that requirement introduces a second channel:

errors := make(chan error, 100)

...

// check errors before using results
select {
case err := <-errors:
    fmt.Println("finished with error:", err.Error())
default:
}

Workers now send to results on success or errors on failure. This breaks the completion signal—you can no longer rely on a closed results channel as the end-of-work marker. A sync.WaitGroup becomes necessary to track when all workers have finished irrespective of outcome:

package main

import (
	"fmt"
	"sync"
	"time"
)

func worker(id int, wg *sync.WaitGroup, jobs <-chan int, results chan<- int, errors chan<- error) {
	for j := range jobs {
		fmt.Println("worker", id, "processing job", j)
		time.Sleep(time.Second)

		if j%2 == 0 {
			results <- j * 2
		} else {
			errors <- fmt.Errorf("error on job %v", j)
		}
		wg.Done()
	}
}

func main() {
	jobs := make(chan int, 100)
	results := make(chan int, 100)
	errors := make(chan error, 100)

	var wg sync.WaitGroup
	for w := 1; w <= 3; w++ {
		go worker(w, &wg, jobs, results, errors)
	}

	for j := 1; j <= 9; j++ {
		jobs <- j
		wg.Add(1)
	}
	close(jobs)

	wg.Wait()

	select {
	case err := <-errors:
		fmt.Println("finished with error:", err.Error())
	default:
	}
}

The code remains serviceable but loses its original crispness. Worse, a subtle bug creeps in.

A Deadlock Waiting to Happen

If the error channel is unbuffered or too small—say, capacity 1—and more than that many jobs fail, workers block trying to enqueue errors. Nobody is reading that channel until the WaitGroup completes, which never happens. Deadlock:

errors := make(chan error, 1)
$ go run worker_pool_err.go
worker 3 processing job 1
worker 1 processing job 2
worker 2 processing job 3
worker 2 processing job 5
worker 1 processing job 4
worker 1 processing job 6
worker 1 processing job 7
fatal error: all goroutines are asleep - deadlock!

This failure mode is fixable, but it proves a point: building a production-grade worker pool requires more thought than the toy examples suggest.

A Sturdier Approach

A more robust design avoids the secondary channel entirely by associating errors with tasks. The task object carries its own error field, eliminating the risk of channel saturation:

import (
	"sync"
)

// Pool is a worker group that runs a number of tasks at a
// configured concurrency.
type Pool struct {
	Tasks []*Task

	concurrency int
	tasksChan   chan *Task
	wg          sync.WaitGroup
}

// NewPool initializes a new pool with the given tasks and
// at the given concurrency.
func NewPool(tasks []*Task, concurrency int) *Pool {
	return &Pool{
		Tasks:       tasks,
		concurrency: concurrency,
		tasksChan:   make(chan *Task),
	}
}

// Run runs all work within the pool and blocks until it's
// finished.
func (p *Pool) Run() {
	for i := 0; i < p.concurrency; i++ {
		go p.work()
	}

	p.wg.Add(len(p.Tasks))
	for _, task := range p.Tasks {
		p.tasksChan <- task
	}

	// all workers return
	close(p.tasksChan)

	p.wg.Wait()
}

// The work loop for any single goroutine.
func (p *Pool) work() {
	for task := range p.tasksChan {
		task.Run(&p.wg)
	}
}

This pool implementation, used to build the author's website, defines its workload through a task interface:

// Task encapsulates a work item that should go in a work
// pool.
type Task struct {
	// Err holds an error that occurred during a task. Its
	// result is only meaningful after Run has been called
	// for the pool that holds it.
	Err error

	f func() error
}

// NewTask initializes a new task based on a given work
// function.
func NewTask(f func() error) *Task {
	return &Task{f: f}
}

// Run runs a Task and does appropriate accounting via a
// given sync.WorkGroup.
func (t *Task) Run(wg *sync.WaitGroup) {
	t.Err = t.f()
	wg.Done()
}

Callers define task types and hand them off. Error handling occurs after all work completes, by inspecting each task's error field:

tasks := []*Task{
    NewTask(func() error { return nil }),
    NewTask(func() error { return nil }),
    NewTask(func() error { return nil }),
}

p := pool.NewPool(tasks, conf.Concurrency)
p.Run()

var numErrors int
for _, task := range p.Tasks {
    if task.Err != nil {
        log.Error(task.Err)
        numErrors++
    }
    if numErrors >= 10 {
        log.Error("Too many errors.")
        break
    }
}

The pattern is straightforward, but each project currently vendors its own variation. It's almost too small to justify an external dependency—yet hundreds of subtly different implementations litter GitHub, each with its own API and edge cases. Community convergence on a single third-party package seems unlikely.

This seems like a case where the Go maintainers could provide real guidance. A worker pool in the standard library would give developers a default, well-tested path, preventing the fragmentation that comes from dozens of independent solutions to the same problem.