Why flag parsing deserves test coverage
A program that parses flags and other command-line arguments generally follows a familiar flow: define a flag set, parse the arguments, then invoke the corresponding behavior for each parsed flag and any remaining positional arguments.
For small tools, that parse step may be trivial enough that unit tests feel unnecessary. In command-line utilities such as grep or find, though, flags are a major part of the user-visible functionality, so testing them matters. Two kinds of tests are relevant here:
- Unit tests covering only the "parse flags" portion of the program.
- Integration tests whose input is the full command line and whose output is the program's total effect — observable output and/or environment changes.
The technique below serves both cases. It relies on standard flag package facilities; projects built on a framework like Cobra should consult that framework's documentation instead.
Separating parsing from execution
To make the parsing step testable, split the flow so that instead of parsing a flag and immediately invoking its behavior, the parsed values are stored in a Config value. That value can then be inspected directly from tests.
A sample parseFlags for a toy program looks like this. The complete code, tests included, accompanies the original post.
type Config struct {
verbose bool
greeting string
level int
// args are the positional (non-flag) command-line arguments.
args []string
}
// parseFlags parses the command-line arguments provided to the program.
// Typically os.Args[0] is provided as 'progname' and os.args[1:] as 'args'.
// Returns the Config in case parsing succeeded, or an error. In any case, the
// output of the flag.Parse is returned in output.
// A special case is usage requests with -h or -help: then the error
// flag.ErrHelp is returned and output will contain the usage message.
func parseFlags(progname string, args []string) (config *Config, output string, err error) {
flags := flag.NewFlagSet(progname, flag.ContinueOnError)
var buf bytes.Buffer
flags.SetOutput(&buf)
var conf Config
flags.BoolVar(&conf.verbose, "verbose", false, "set verbosity")
flags.StringVar(&conf.greeting, "greeting", "", "set greeting")
flags.IntVar(&conf.level, "level", 0, "set level")
err = flags.Parse(args)
if err != nil {
return nil, buf.String(), err
}
conf.args = flags.Args()
return &conf, buf.String(), nil
}
Two features of the flag package make this work:
- Creating a custom
flag.FlagSetrather than using the default global one. - Using the
XxxVarvariants of the flag definition methods (for exampleBoolVar) so parsed values are written into pre-defined variables.
Handling usage flags such as -h requires a little care; apart from that the code is straightforward. The remainder of the program simulates arbitrary work based on the parsed configuration.
func doWork(config *Config) {
fmt.Printf("config = %+v\n", *config)
}
func main() {
conf, output, err := parseFlags(os.Args[0], os.Args[1:])
if err == flag.ErrHelp {
fmt.Println(output)
os.Exit(2)
} else if err != nil {
fmt.Println("got error:", err)
fmt.Println("output:\n", output)
os.Exit(1)
}
doWork(conf)
}
Unit tests and integration tests
Once parsing is isolated, unit-testing the flags is simple.
func TestParseFlagsCorrect(t *testing.T) {
var tests = []struct {
args []string
conf Config
}{
{[]string{"-verbose"},
Config{verbose: true, greeting: "", level: 0, args: []string{}}},
{[]string{"-level", "8", "-greeting", "joe", "-verbose", "foo"},
Config{verbose: true, greeting: "joe", level: 8, args: []string{"foo"}}},
// ... many more test entries here
}
for _, tt := range tests {
t.Run(strings.Join(tt.args, " "), func(t *testing.T) {
conf, output, err := parseFlags("prog", tt.args)
if err != nil {
t.Errorf("err got %v, want nil", err)
}
if output != "" {
t.Errorf("output got %q, want empty", output)
}
if !reflect.DeepEqual(*conf, tt.conf) {
t.Errorf("conf got %+v, want %+v", *conf, tt.conf)
}
})
}
}
An integration test driven by a command line looks much the same. The difference is the comparison target: instead of checking a Config against an expected value, you compare the output and/or side effects of doWork. Such tests are typically split into several table-driven variants reflecting the nature of doWork. If its only observable side effect is text output, this stays simple. If it touches the file system or runs network servers or clients, things get trickier — but no different from testing any program with those behaviors, flag parsing aside.
The accompanying code includes a fuller set of tests that exercise error scenarios as well.
| [1] | Given the invocation grep -i joe file.go, we'll refer to -i as a flag and to joe and file.go as positional arguments. The distinction is not terribly important for the sake of this post, however. |



