API reference
Every exported name, grouped by the job it does. Each docstring opens with the situation it serves. The 0.4 spellings that still work, with a warning, are listed last; Migration from 0.4 shows how to rewrite them and what was removed.
UnitTestDesign.UnitTestDesign — Module
Use when a function or a configuration has several options and you want to test their combinations: describe the configurations your code must handle; it tells you which combinations your tests exercise, and supplies a compact set of additional cases covering the rest.
UnitTestDesignStart with TestSpace and all_pairs, and measure any set of cases with coverage.
Spaces and values
A TestSpace names the parameters of a test and lists their values. Values are kept as given; Invalid marks a value for negative tests and Partition stands for a class of values drawn at run time.
UnitTestDesign.TestSpace — Type
Use when you want to describe a test's parameters, their values, and the rules that exclude combinations once, and share that description among generation, measurement and diagnosis.
TestSpace(domains::NamedTuple; constraints = [], tabulation_limit = 10^5)
TestSpace(name => domain, ...; constraints = [], tabulation_limit = 10^5)The parameters of a test, their values, and the rules that exclude combinations (contract §2, §12.11). Build one when a function or a configuration has several options and not every combination is valid:
space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
])Parameter names are distinct Symbols, at least one, and none is nothing or missing (§2.10, §12.7). Each domain is a nonempty vector, range, or tuple, in the order its values should be tried (§2.6); a single value is allowed (§2.7). Values keep their identity: the same concrete type and isequal, so Any[1, 1.0] is two choices and no value is converted (§2.1–§2.3). A domain listing the same choice twice is an error (§2.5). nothing and missing are ordinary values (§2.9). Domains may hold Partition and Invalid values (§4, §5); each parameter needs an ordinary value.
constraints are rules built with forbid, require, @forbid and @require. The space checks each rule's names and pattern values, then tabulates it: a rule is evaluated once per combination of its parameters' values, when the space is built, unless there are more than tabulation_limit combinations, in which case it is evaluated lazily and the package warns (§12.18–§12.20).
Domains are copied, preserving their element types; tuples become vectors. Values themselves are not copied (§2.8).
parameters(space) gives the names and length(space) the size of the full product, including rows the rules exclude.
UnitTestDesign.parameters — Function
Use when you need a space's parameter names in order, for example to name the columns of a positional result: DataFrame(cases, parameters(cases.space)).
parameters(space::TestSpace) -> Vector{Symbol}The parameter names, in order.
UnitTestDesign.Invalid — Type
Use when a parameter has values the code must reject, and you want negative cases that try each invalid value beside otherwise valid values, one invalid value per case.
Invalid(x)Marks x as an invalid value of the parameter whose domain lists it, for negative tests (contract §5.1). A row with one invalid value is a negative row: it must satisfy only the rules whose scope omits that parameter (§5.5). Rules never receive an Invalid (§5.8), and every parameter needs at least one ordinary value (§5.2).
Invalid(x) and Invalid(y) are the same choice exactly when x and y are; Invalid(x) and x are different choices, so a domain may hold both (§2.12). Invalid(Invalid(x)) and Invalid(Partition(...)) are errors (§4.12).
Generation builds the covering design over ordinary values, then adds negative rows: each holds one invalid value beside ordinary values, and together they cover that value's negative targets (§6), such as, at strength 2, the invalid value beside every feasible value of every other parameter. Rows keep the wrapper, so a test body can branch on hasinvalid.
The wrapped value is x.value, the supported way to read it. A test body unwraps a row's values before the call with unwrap(x) = x isa Invalid ? x.value : x.
UnitTestDesign.Partition — Type
Use when one value stands for a class of inputs, such as tiny tolerances, and each run should draw a concrete member of the class while rules and coverage count the class by its name.
Partition(name::Symbol, draw)A named choice that stands for a class of values (contract §4.1). Rules, patterns and coverage see name (§4.5); returned rows keep the wrapper, and realize(case; rng) replaces it by draw(rng), a concrete value drawn at run time (§4.7). A fixed value is Partition(:tiny, Returns(1e-9)).
A partition's identity is its name; draw is not part of it (§2.13). Names are unique within a parameter, and a domain may not also hold the raw Symbol of a partition's name (§4.3, §4.4). A partition is an ordinary value (§4.2). Invalid(Partition(...)) is an error (§4.12).
A partition prints as Partition(:tiny), which does not read back as code. In cases committed to a test file as a literal, such as the output of repr(collect(cases)), write its name, :tiny, which coverage and must_include accept in its place (§2.11).
UnitTestDesign.hasinvalid — Function
Use when a test body must tell a negative case, one that holds an Invalid value, from an ordinary one.
hasinvalid(case) -> BoolTrue when the row (a NamedTuple, Tuple, or vector) holds an Invalid value, so a test body can branch on negative rows (contract §5.1).
UnitTestDesign.realize — Function
Use when cases hold Partition values and the test needs concrete inputs: it draws a value for each partition from the rng you pass.
realize(case; rng) -> case
realize(cases; rng) -> VectorReplace each Partition in a row by a value drawn from it: draw(rng), one call per partition, in parameter order (contract §4.7). Every other value, Invalid markers included, passes through unchanged, and the row keeps its shape: a Tuple stays a Tuple of the same length, and a NamedTuple keeps its names and their order (§4.8).
realize(cases; rng) realizes each row of a vector of rows, such as a TestCases, in order with the same rng, and returns a plain Vector of the realized rows, not a TestCases (§4.9).
rng is required: no draw uses the global random number generator. For a given row and rng state the result is the same, so a seeded generator repeats it (§4.10):
julia> space = TestSpace((tol = [Partition(:tiny, rng -> 1e-9 * rand(rng)), 1e-3], method = [:lu, :qr]));
julia> cases = all_pairs(space)
4 cases (minimal) · strength 2 · Auto: Construction() · 2 parameters · 4 combinations, 4 valid
tol method
1 Partition(:tiny) :lu
2 0.001 :qr
3 0.001 :lu
4 Partition(:tiny) :qr
julia> realize(cases; rng = Xoshiro(1)) == realize(cases; rng = Xoshiro(1))
true
julia> realize((tol = Partition(:big, Returns(1e6)), method = :qr); rng = Xoshiro(1))
(tol = 1.0e6, method = :qr)Realized values carry no coverage claim (§4.11): measure coverage on the labeled rows, and keep them beside the realized inputs. A draw that returns a Partition or an Invalid is an ArgumentError naming the partition (§4.12).
Rules
Rules exclude combinations of values. All four surface forms build the same Constraint, which a space takes in its constraints vector.
UnitTestDesign.forbid — Function
Use when some combinations of values are not valid and you can say which: an exact pattern of values, or a predicate over named parameters that returns true for the forbidden combinations.
forbid(pattern::NamedTuple; reason = nothing)Forbid one exact combination of values, such as forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"). A row is excluded when its value at each named parameter is the same choice as the pattern's value: the same type and isequal, so 1 does not match 1.0 (contract §2.1, §12.4). A Partition may be written as its wrapper or its name. Pattern values must be domain values; the space checks them when it is built. A pattern cannot name an Invalid value (§5.8). There is no require pattern form.
forbid(f, names::Symbol...; reason = nothing)
forbid(names::Symbol...; reason = nothing) do values... endForbid the combinations of the listed parameters for which f returns true. f receives their values positionally, in the order listed (§12.5): forbid(:mode, :solver) do m, s; m == :fast && s != :none end. The names are distinct. A Partition is passed as its name; an Invalid value is never passed (§5.8, §12.14).
forbid(f; reason = nothing)
forbid(; reason = nothing) do case; ... endWith no names, a whole-case rule: f receives the complete row as a NamedTuple (§12.9). Use it as an escape hatch. A whole-case rule is evaluated lazily, row by row (§12.20); it connects every parameter into one component, so deciding feasibility may search up to the product of the unassigned domains, bounded by feasibility_limit (§12.21). A whole-case rule does not apply to a row with an Invalid value (§5.6).
Every predicate must return a Bool (§12.15) and should be deterministic and free of side effects: it may be called more than once (§12.17). The rule's reason labels it in errors and reports (§12.3).
UnitTestDesign.require — Function
Use when it is easier to say which combinations are valid than which are not: the rule excludes every combination for which the predicate returns false.
require(f, names::Symbol...; reason = nothing)
require(names::Symbol...; reason = nothing) do values... end
require(f; reason = nothing)Allow only the combinations for which f returns true; a row where it returns false is excluded (contract §12.2). The forms and the arguments f receives are those of forbid: listed names receive their values positionally, and no names means a whole-case rule that receives the complete row as a NamedTuple. require(:rows, :cols) do r, c; r == c end allows only square shapes. There is no require pattern form (§12.4).
A require rule is stored as the forbid rule "f is false", with its polarity kept for display.
UnitTestDesign.@forbid — Macro
Use when a rule reads most clearly as a Julia expression over bare parameter names that is true for the forbidden combinations, such as @forbid mode == :fast && solver != :none.
@forbid expr
@forbid(expr; reason = "...")Forbid the combinations for which expr is true, written with bare parameter names: @forbid mode == :fast && solver != :none.
Which identifiers are parameters (contract §12.6, §12.7):
- Every free identifier that is not in call position names a parameter. The rule's scope lists them in order of first appearance.
- An identifier in call position is an ordinary function (
isodd(n),n < 3), as is a dotted name (Base.isodd(n),M.x). To read a field of a parameter, writegetproperty(p, :field). nothing,missing,true,false, and literals (:fast,1e-3,"s",r"re") are values, so no parameter may be namednothingormissing.$xinterpolates the caller'sx, evaluated once when the rule is built:@forbid n < $threshold. Write$Inf,$Intand the like for any other global that is not called.- Names bound inside the expression are local, not parameters, with Julia's scoping: the arguments of
->and of an anonymousfunction(including keyword arguments anddo-block arguments),letbindings, and the variables of generators and comprehensions. So in@forbid (n -> n)(m) > nthe scope is(m, n): the lambda'snis local. - Anything else that binds or assigns a name or runs statements (an assignment outside a
letbinding,for,while,try,global,local, a quoted expression, or a macro call) is anArgumentErrorwhen the macro expands. Write such a rule with the function form,forbid(f, names...). - A subtype test written with the operator,
T <: $AbstractFloatorT >: $Int, is anArgumentErrortoo:<:and>:are syntax, not calls. Write the call,(<:)(T, $AbstractFloat), or use the function form.
The rule's label is its source text as Julia prints it, such as "@forbid(mode == :fast && solver != :none)", preceded by the reason if one is given (§12.3). A name the space lacks is an error when the space is built, suggesting $name if a variable was meant (§12.8). The macro builds the same Constraint as forbid with listed names.
UnitTestDesign.@require — Macro
Use when a rule reads most clearly as a Julia expression over bare parameter names that must be true in every valid combination, such as @require mode == :exact || solver == :none.
@require expr
@require(expr; reason = "...")Allow only the combinations for which expr is true, written with bare parameter names: @require mode == :exact || solver == :none. Identifiers are read, and binding forms accepted or rejected, as in @forbid. The macro builds the same Constraint as require with listed names.
UnitTestDesign.Constraint — Type
Use when you handle rules as values: forbid, require, @forbid and @require all return a Constraint, and a TestSpace takes a vector of them as constraints.
ConstraintOne rule of a TestSpace: the single internal form that forbid, require, @forbid and @require all build (contract §12.1). Nothing downstream of construction distinguishes the surface forms; a space tabulates every rule the same way.
Fields:
scope::Tuple{Vararg{Symbol}}: the parameters the rule reads, in the order its predicate receives them. The empty tuple marks a whole-case rule, whose scope is every parameter (§12.9).predicate: returnstruewhen the combination is forbidden. A scoped rule's predicate takes the scoped values positionally, in scope order (§12.5); a whole-case rule's takes the complete row as oneNamedTuple. For arequirerule it is the negation of the caller's function. A result that is not aBoolpasses through unchanged so that evaluation can report it (§12.15).polarity::Symbol::forbidor:require. Display only (§12.2).label::String: thereason, the macro's source text, or both. Empty when the rule has neither; the space then names the rule by its position and scope (§12.3, seerule_label).source::Symbol::pattern,:names,:macro, or:whole_case. Used only to phrase construction errors (the$namehint of §12.8).pattern: for a pattern rule, the pattern as given, so the space can check its values against the domains (§12.4); otherwisenothing.
show prints the label.
UnitTestDesign.ConstraintError — Type
Use when a rule's predicate may throw: the exception reaches you wrapped in a ConstraintError that names the rule and the values it received.
ConstraintError(rule, arguments, exception)A rule's predicate threw an exception (contract §12.16). rule names the rule by position and label, arguments is a NamedTuple of the values it received (for a whole-case rule, the row), and exception is what it threw, which is also the cause on the exception stack. An exception never means forbidden or allowed: fix the predicate so that it returns true or false for every combination of its parameters' values.
Generation
covering is the general entry point; all_values, all_pairs and all_triples are covering at strength 1, 2 and 3. excursions and full_factorial are the other two strategies. Every generator returns a TestCases.
UnitTestDesign.covering — Function
Use when you want every combination of values of every strength parameters (pairs at strength 2, triples at 3) to appear in at least one case, with as few cases as the engine finds.
covering(space; strength = 2, stronger = [], must_include = [], engine = Auto(),
feasibility_limit = 1_000_000, explanation_limit = 1_000_000)
covering(domains::NamedTuple; constraints = [], kwargs...)
covering(name => domain, ...; constraints = [], kwargs...)
covering(domain, domain, ...; kwargs...)Test cases in which every combination of values of every strength parameters appears at least once, among the combinations some valid row contains; combinations the rules exclude need no case (contract §1.2, §1.3). Returns a TestCases, a vector of rows that also records the request and what it excluded (§1.18, §1.19).
The parameters come in one of four forms:
- a
TestSpace, which holds the parameters, their values and the rules.constraints =is an error here: rules belong to the space (§12.12). - a
NamedTupleof domains,(mode = [:fast, :exact], tol = [1e-3, 1e-6]), orname => domainpairs,:mode => [:fast, :exact], :tol => [1e-3, 1e-6].constraints =builds the space, asTestSpace(domains; constraints)would. Rows areNamedTuples. - one domain per parameter, a vector, range or tuple each:
covering([1, 2, 3], ["a", "b"], [true, false]). The parameters are namedp1,p2, ... in messages, and rows are tuples. Positional calls take noconstraints; to exclude combinations, name the parameters.
Values are kept as given, with their types: Any[1, 1.0] is two values, and nothing and missing are ordinary values (§2.1, §2.9).
A Partition is an ordinary value that rules and targets see by its name; returned rows hold the wrapper, and realize draws the concrete values (§4). An Invalid value is for negative tests (§5, §6). The covering design is built over ordinary values; then, for each invalid value v of a parameter p, negative rows hold p = v beside ordinary values of the other parameters, and cover every feasible combination of v with strength - 1 other parameters' values, and within each stronger group that contains p, with its strength less one. At strength 2, each invalid value appears beside every value of every other parameter that some valid negative row holds (§6.6). A negative row satisfies the rules that do not read p; rules that read p do not apply to it (§5.5). The rows are the must-include rows, then the ordinary rows, then the negative rows (§5.12), and a row never holds two Invalid values (§5.7). hasinvalid(case) tells a test body which kind it has, and the result counts the two kinds of targets separately (§5.10).
Keywords
strength = 2: from 1 to the number of parameters (§11.1, §11.2). At the number of parameters the result is, as a set, every valid row (§7.8).stronger = []: groups of parameters that need a higher strength, asgroup => strengthpairs. A named group is a tuple of names,[(:a, :b, :c) => 3]; a positional group lists argument positions,[(1, 3, 4) => 3]. Overlapping groups combine; a group at the base strength adds nothing. The caller's vector is not changed (§11.3–§11.9).must_include = []: rows that must appear. They come first, in the order given, duplicates kept (§10.5). A named call takesNamedTuples, which may be partial and are then completed with valid values, or an existingTestCases, whose rows are kept and topped up with the rows needed to cover what they miss (§9.10). A positional call takes tuples or vectors with one value per parameter. A row that breaks a rule, or a partial row with no valid completion, is an error naming the row and the rules (§10.2–§10.4). A row with oneInvalidvalue is a negative row, judged and completed under the negative policy, with ordinary values elsewhere; a partial row without one is completed as an ordinary row; a row with two is an error (§5.7, §7.9).engine = Auto(): the covering engine.Auto(), the default, keeps the smaller of IPOG's design and the catalog's array (Construction) where the catalog applies, and gives IPOG's design elsewhere;recommendsays what it would run.IPOG()alone is fast and covers any request;Auto(goal = :compact)then removes rows with the row reducer (Compact);GNDis a seeded greedy search. Every engine is deterministic for the same inputs (§9.1). None guarantees a particular number of cases, and none is always smaller than another (§8.1); the result records a lower bound beside its count, and says "minimal" when the count meets it (§8.4).feasibility_limit = 1_000_000: the node budget of each search that decides whether a combination, a partial must-include row or a placement has a valid completion (§3.3, §3.4). Generation resolves every one of them or fails: a search that runs out throws aResourceLimitErrornamingfeasibility_limit, and no design is returned with an undecided combination (§1.20, §3.6, §3.7). Raise it when generation cannot finish, as inall_pairs(space; feasibility_limit = 10_000_000); a larger value never changes the rows of a result that already succeeded, though an implied exclusion's explanation may differ (§3.8).explanation_limit = 1_000_000: the node budget of the search that finds which rules cause each implied exclusion, starting from a set of rules already proven to exclude it (§3.13, §3.14). Each trial of that search is also bounded byfeasibility_limit. Running out never throws and never changes the cases: the design is complete and certified as usual, and the exclusion keeps its proven set of rules withminimal = :unresolved, whichshowcounts as "with an unresolved explanation" (§3.15, §3.16). Itslimitis:explanation_limit => Nwhen this budget stopped a trial or left one untried, and:feasibility_limit => Nwhen a trial reached the feasibility limit instead (§3.14). Raise it only when you want each implied exclusion's rules verified inclusion-minimal, so that none of them can be dropped; only the attribution becomes more precise. It also bounds the explanation in the error for a partial must-include row with no valid completion.
The 0.4 keywords n_way (now strength), seeds (now must_include) and wayness (now stronger, translated from its Dict of positions) are accepted with a deprecation warning (§13.1). Passing one together with its new keyword is an ArgumentError, even when the new keyword is at its default value, as in strength = 2, n_way = 1 (§13.2).
Examples
space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [@require(mode == :exact || solver == :none)])
covering(space; strength = 2)
covering((a = 1:3, b = [:x, :y], c = [true, false]); strength = 3)
covering(fill(1:4, 10)...; stronger = [(1, 2, 3) => 3], engine = GND())
all_pairs(space; must_include = previous_cases) # keep them; add what they missSee also all_values, all_pairs, all_triples, excursions, full_factorial.
UnitTestDesign.all_values — Function
Use when you want every value of every parameter to appear in at least one case, with as few cases as the engine finds: a quick check that each value works at all.
all_values(input...; stronger, must_include, engine, constraints,
feasibility_limit, explanation_limit)Test cases in which every value of every parameter appears at least once: covering at strength 1. It takes the same inputs and every keyword of covering except strength.
all_values([1, 2, 3], ["a", "b"], [true, false]) # 3 tuplesUnitTestDesign.all_pairs — Function
Use when you want every pair of parameter values to appear in at least one case, with as few cases as the engine finds.
all_pairs(input...; stronger, must_include, engine, constraints,
feasibility_limit, explanation_limit)Test cases in which every pair of values of every two parameters appears at least once, among the pairs some valid row contains: covering at strength 2. It takes the same inputs and every keyword of covering except strength.
all_pairs([1, 2, 3], ["a", "b", "c"], [true, false])
all_pairs((mode = [:fast, :exact], solver = [:none, :lu, :qr]);
constraints = [@forbid(mode == :fast && solver != :none)])
all_pairs(space; must_include = existing_cases)UnitTestDesign.all_triples — Function
Use when you want every combination of three parameters' values to appear in at least one case, with as few cases as the engine finds; it reaches faults that need three values together, at the cost of more cases than pairs.
all_triples(input...; stronger, must_include, engine, constraints,
feasibility_limit, explanation_limit)Test cases in which every combination of values of every three parameters appears at least once, among those some valid row contains: covering at strength 3. It takes the same inputs and every keyword of covering except strength.
all_triples([1, 2], [3, 4], [5, 6], [7, 8])UnitTestDesign.excursions — Function
Use when you trust one base case and want every valid variation that changes at most distance of its parameters. An excursion is not a covering design: it does not guarantee that every pair, or even every value, appears.
excursions(space; from = nothing, distance = 1, must_include = [],
feasibility_limit = 1_000_000, explanation_limit = 1_000_000)
excursions(domains::NamedTuple; constraints = [], kwargs...)
excursions(name => domain, ...; constraints = [], kwargs...)
excursions(domain, domain, ...; kwargs...)Variations around one base row: the must-include rows, in the order given with duplicates kept (§10.5), then the base, then every other valid row that differs from the base in at most distance parameters (contract §7.5). The base and each of those rows appear once, and a row equal to a must-include row is not repeated (§7.11). Returns a TestCases with strategy :excursion. The inputs are the four forms covering takes.
The promise is only this: every returned row other than a must-include row is within distance changed parameters of the base. An excursion is not a covering design. It makes no claim that every pair, or even every value, appears: a value whose rows within the distance all break a rule appears in no row (§7.7).
from: the base, a complete valid ordinary row, as aNamedTupleor, for a positional call, a tuple of values in argument order. Omitted, it is the first ordinary value of each parameter. A base that is partial, holds anInvalidvalue, or breaks a rule is an error naming the cause (§7.6). The base is never dropped.distance = 1: an integer of at least 0. 0 gives the must-include rows and the base alone; a distance above the number of parameters is the number of parameters. Excursion distance is not covering strength, and there are nostrongergroups (§7.5).must_include: rows placed first, as forcovering(§10), and kept as given, duplicates included. A partial row is completed toward the base. An excursion row equal to a must-include row is not repeated (§7.11).
Changing a parameter to one of its Invalid values gives a negative row, which is kept when it satisfies the rules that do not read that parameter (§5.5, §7.5). A row with two Invalid values is never returned (§5.7).
Rows within the distance that break a rule are left out. The result reports how many in cases.notes.dropped, and cases.notes.never_appear lists the values that appear in no returned row (§7.7).
feasibility_limit and explanation_limit bound only the searches for must-include rows, as for full_factorial.
excursions([1, 2, 3], [:x, :y], [true, false]) # 1 + 2 + 1 + 1 rows
excursions(space; from = (mode = :exact, solver = :lu, tol = 1e-6), distance = 2)UnitTestDesign.full_factorial — Function
Use when the product of the domains is small enough to run every valid combination, or when you need every one of them.
full_factorial(space; limit = 10^6, must_include = [], feasibility_limit = 1_000_000)
full_factorial(domains::NamedTuple; constraints = [], kwargs...)
full_factorial(name => domain, ...; constraints = [], kwargs...)
full_factorial(domain, domain, ...; kwargs...)Every valid row: the full product of the domains, less the rows the rules exclude (contract §7.2). The must-include rows come first, in the order given with duplicates kept (§10.5); then each remaining valid row appears once, and a valid row equal to a must-include row is not repeated. Returns a TestCases with strategy :full_factorial; the inputs are the four forms covering takes.
With Invalid values, the valid ordinary rows come first, then the valid negative rows: for each parameter in order and each of its invalid values in domain order, the rows holding that value beside ordinary values of the other parameters, kept when they satisfy the rules that do not read that parameter (§5.5). A row with two Invalid values is never returned (§5.7).
limit guards against a product too large to enumerate. Before looking at any row, must-include rows included, the call counts the candidate rows, the product of the parameters' ordinary value counts plus, for each parameter, its number of Invalid values times the product of the other parameters' ordinary value counts, and if that exceeds limit it throws a ResourceLimitError that gives the count and the keyword (§7.3), so no search runs for a refused enumeration. Raise limit to go ahead, or use covering for a smaller design. Candidates are then enumerated one at a time and only valid rows are kept (§7.4). cases.notes reports candidates and accepted separately.
The enumeration checks complete rows and does not search, so feasibility_limit and explanation_limit matter only for a partial must-include row. feasibility_limit bounds the search that completes it: running out throws a ResourceLimitError naming the keyword, and raising it is how to let the call finish. explanation_limit bounds the search that explains a row with no valid completion: running out still throws the ArgumentError for that row, naming rules proven to exclude it, with the explanation marked unresolved; raising it only narrows that list of rules.
full_factorial([0.1, 0.2, 0.3], ["low", "high"], [false, true]) # 12 tuples
full_factorial((a = 1:3, b = [7, 8], c = [true, false]);
constraints = [forbid((b = 7, c = false))]) # 9 rowsUnitTestDesign.TestCases — Type
Use when you work with what a generator returned: a read-only vector of cases that also records the request, the engine and seed, and the combinations the rules excluded.
TestCases{T} <: AbstractVector{T}The cases a generation call returns. T is a NamedTuple type for named spaces and a Tuple type for positional calls (§1.18). Each field's type comes from the parameter's values, not from the domain's element type: the one concrete type they share, or else the Union of their concrete types (§2.4). So [1, 2, 3] and Any[1, 2] give an Int field, Any[1, 1.0] a Union{Int64, Float64} field, [nothing, :x] a Union{Nothing, Symbol} field, and [1, Invalid(1)] a Union{Int64, Invalid{Int64}} field. No value is converted to fit its field: 1 stays an Int and 1.0 a Float64. Rows are immutable.
A vector of rows
A TestCases is a read-only AbstractVector{T}: length, cases[i], first, last, eachindex and iteration work as for any vector, and for (mode, solver, tol) in cases destructures each row. collect(cases) and copy(cases) are plain, mutable Vector{T}s, and so is a slice such as cases[1:2] or cases[[1, 3]]; the bookkeeping stays with the original. == compares rows, as for any vector, so cases == collect(cases).
Fields
All recorded at generation and never recomputed (§1.19, §1.22): cases, space, strategy (:covering, :excursion, :full_factorial), strength (the covering strength, see below), stronger (as names => strength pairs, base group excluded), engine::Symbol, seed, n_must_include, required and covered (ordinary target counts; zero for non-covering strategies), excluded (ordinary Exclusions, in target order, §9.7), positional::Bool, notes, which is strategy specific (§7.3, §7.7), and the negative bookkeeping, kept apart from the ordinary (§5.10, §6): negative_required and negative_covered (negative target counts, zero without Invalid values and for non-covering strategies) and negative_excluded (the negative targets no valid negative row can hold, as Exclusions in the order coverage lists them); and record, below. notes holds:
- an excursion's
base, the base row, of the result's row typeT;distance, after clamping to the parameter count;dropped, the number of rows within the distance that broke a rule; andnever_appear, the values that appear in no returned row, as aVector{Pair{Symbol, Any}}ofname => valuein parameter and domain order (:p2 => 3for a positional result); - a full factorial's
candidates(the ordinary product plus the rows with oneInvalidvalue, §7.3) andaccepted(the valid rows, ordinary and negative); - nothing for a covering design.
strength is 0 when the strategy has no strength (excursions and full factorials); measurement of such a result needs an explicit strength.
The record
record is a NamedTuple of plain data on how the cases were made (plan §4.1, §4.2, §6.1). Its first fields are the package's own, computed and checked by generation, never reported by an engine:
randomized: whether the engine drew random numbers, so thatseedrepeats the cases (§9.5).lower_bound: for a covering design, a proven lower bound on the cases of any design for the same request, andproof, why, as "the 3 × 3 = 9 combinations of a and b need a case each": every case holds one combination of each set of parameters, so one set's required combinations need as many cases; the must-include rows are cases of every such design; and the negative rows of eachInvalidvalue are bounded in the same way and added.minimalistruewhen the cases number exactlylower_bound: then no design for the request has fewer, and this is the one case in which the package calls a count minimal (contract §8.4).nothing,falseand""for an excursion or a full factorial.engine: the engine's configuration, the whole tree of it:name;call, its constructor call, which with the same request repeats the cases unless an engine in it drew from a caller'srng;seed;randomized; andsettings, its other settings, where an engine it wraps (innerofCompact) or chooses among (candidatesofAuto) is a configuration of the same form. ForCompact(GND(seed = 17); seed = 3, effort = 2):(name = :Compact, call = "Compact(GND(seed = 17); seed = 3, effort = 2)", seed = 3, randomized = true, settings = (inner = (name = :GND, call = "GND(seed = 17)", seed = 17, randomized = true, settings = (candidates = 50,)), effort = 2))
Then the stages that ran, each a NamedTuple whose engine is the call of the engine that ran and rows the rows it made, followed by what that engine reports:
ordinary: the ordinary design's stage.Autoreportschose, the engine call that made the ordinary cases, such as"Construction()"or"Compact(IPOG())";starts, the stage of each start it ran;kept, the index of the one kept among them; and withgoal = :compactreducer. A catalog array (Construction) reportscatalog, with the construction'sname,family,source,rows, the array'slower_bound, whether the design is anorthogonalarray (every combination ofstrengthparameters exactly once, so never withInvalidvalues, whose negative rows repeat ordinary combinations), and whether the array onlyseededthe design under rules. The row reducer (Compact) reportsstart, the stage of its inner engine, andreducer, with the rows it started from and ended with, its bound, its steps and budgets, and why it stopped.IPOGreportsmember, the member of the IPOG family whose design it kept,(tiebreak, vertical), or(tiebreak = :none, vertical = :none)at full strength, where the design is every valid row and no member runs (the manual's IPOG page says what the members are).GNDreports nothing more. ForAuto()on eight parameters of seven values:(engine = "Auto()", rows = 49, chose = "Construction()", starts = [(engine = "Construction()", rows = 49, catalog = (name = "Bush", family = "Bush orthogonal array", …))], kept = 1)negative: for eachInvalidvalue, in parameter and domain order,(parameter, value, rows, stage): the parameter's name, the value as itsrepr, the cases that hold it, and the stage that covered its negative targets, a request one strength lower on the other parameters, whoseenginesays which engine ran: the same one, or IPOG where the engine has nothing for that request, asConstructionat strength 1.stageisnothingwhere the value needs no such request (strength 1, with nostrongergroup holding the parameter), andnegativeis empty withoutInvalidvalues.
An excursion and a full factorial have no engine: engine, ordinary and negative are nothing.
Display
show prints a summary line, the excluded targets' counts, and the rows as a table (§1.22):
5 cases (lower bound 4) · strength 2 · Auto: IPOG() · 3 parameters · 12 combinations
excluded: 3 pairs forbidden, 2 impossible under the constraints; see report(cases)
mode solver tol
1 :fast :none 0.001
2 :exact :none 1.0e-6
3 :exact :lu 1.0e-6
4 :exact :qr 1.0e-6
5 :fast :none 1.0e-6The summary gives the count, with a covering design's recorded lower bound beside it, as "(lower bound 4)", or "(minimal)" when the count equals it; then the strategy (the strength, any stronger groups and the engine, with a randomized engine's seed and what Auto chose; an excursion's distance, base and dropped rows; a full factorial), the parameter count and the size of the full product. It adds the number of valid rows only when generation already knows it: for a full factorial, and for a covering design at strength equal to the parameter count, where the required targets are the valid rows. Values print with show, so :fast and "fast" differ.
With Invalid values, a covering design's summary ends with its count of negative targets, the excluded line counts negative exclusions after "negative:", and a negative row, one holding an Invalid value, is marked with ! after its row number. For all_pairs(TestSpace((n = [1, 2, Invalid(1)], m = [:a, :b], k = [:x, :y]); constraints = [@forbid(n == 2 && m == :b), @forbid(m == :a && k == :y)])):
6 cases (lower bound 5) · strength 2 · Auto: IPOG() · 3 parameters · 12 combinations · 4 negative targets
excluded: 2 pairs forbidden, 1 impossible under the constraints; see report(cases)
n m k
1 1 :a :x
2 1 :b :y
3 2 :a :x
4 1 :b :x
5! Invalid(1) :a :x
6! Invalid(1) :b :yIn a REPL (an IOContext with :limit => true), long results keep their first 10 and last 5 rows and wide cells and columns are cut to the display size, as a DataFrame does; otherwise every row prints in full. Inside a container, a TestCases prints as its summary line. Display performs no search, rule evaluation or coverage count; verification and bonus coverage belong to report and coverage (§1.23).
Tables
A TestCases of named rows is a vector of NamedTuples, which Tables.jl reads as a row table: DataFrame(cases) has one column per parameter, with the field types above, and CSV.write(path, cases) writes a header of parameter names and one line per case. CSV is text: a Symbol is written as its name and reads back as a string (:fast becomes "fast"), and 1 and 1.0 in a Union{Int64, Float64} column are written as 1 and 1.0. CSV.jl refuses a nothing value; write CSV.write(path, cases; transform = (column, value) -> something(value, missing)) to write it as an empty field, which reads back as missing.
A positional result is a vector of Tuples, which Tables.jl does not recognize as a table (Tables.istable(cases) is false): DataFrame(cases) still builds, with columns named 1, 2, 3, and CSV.write fails. Name the columns yourself:
DataFrame(cases, parameters(cases.space)) # columns p1, p2, p3
CSV.write(path, NamedTuple{Tuple(parameters(cases.space))}.(cases))UnitTestDesign.Exclusion — Type
Use when you want to know why a design holds no case with some combination: an Exclusion names a target no valid case can hold and the rules that exclude it.
ExclusionOne target a design did not need to cover, in the user's vocabulary (contract §1.4): target (a partial NamedTuple), status (:forbidden or :implied), rules (constraint positions in the space), labels (their display labels), minimal (:verified, :unresolved, :not_applicable), and limit (keyword => value when a limit cut the explanation short, else nothing).
It prints as one line naming the target and its rules:
(mode = :exact, tol = 0.001): forbidden by rule 2 (exact mode needs a tight tolerance)
(solver = :lu, tol = 0.001): impossible because rules 1 and 2 combine (rule 1: …; rule 2: …)An unresolved explanation adds the limit that left it unresolved, such as "(explanation unresolved: explanation_limit = 1 reached)"; its rules are still sufficient to exclude the target (§3.15).
UnitTestDesign.ResourceLimitError — Type
Use when a call may stop at a search or size limit: catch this error, or retry with the keyword it names set higher; it never carries a partial result.
ResourceLimitError(what, limit, keyword)A search or enumeration stopped at a resource limit before reaching a conclusion (contract §3.7). what names the operation and the assignment or target being resolved, limit is the value that was reached, and keyword is the keyword that raises it, such as :feasibility_limit. The error never carries a partial result. Retry the call with a larger value of keyword (§3.8); raising a limit never changes an answer that was already resolved.
Engines
The covering functions take an engine. The default, Auto(), chooses among the others by the space and the goal, and recommend says what it would choose before generating; IPOG() alone is Auto(goal = :fast). The Engines page compares them.
UnitTestDesign.IPOG — Type
Use when you want IPOG's design itself, which Auto(goal = :fast) also runs: deterministic and free of randomness, so the same inputs always give the same cases, and covering any request.
IPOG()In-parameter-order General (IPOG): deterministic, no randomness (contract §9.4). It builds the design one parameter at a time, as four members of the IPOG family, and keeps the smallest (the manual's IPOG page). The default engine, Auto(), gives IPOG's design except where the catalog's array (Construction) is smaller.
Lei, Yu, Raghu Kacker, D. Richard Kuhn, Vadim Okun, and James Lawrence. 2008. "IPOG/IPOG-D: Efficient Test Generation for Multi-Way Combinatorial Testing." Software Testing, Verification & Reliability 18 (3): 125-48.
UnitTestDesign.Auto — Type
Use when you want the package to choose the engine for your space: the smallest of the designs it can build quickly, or, with goal = :compact, that design reduced further; recommend shows the choice before you run it.
Auto(; goal = :balanced, seed = 0, effort = 1)A covering engine that picks its method from the request (plan §6.1). goal says what your tests cost, which the space can't tell:
goal | Runs | For |
|---|---|---|
:fast | IPOG alone, the same rows as engine = IPOG() | cheap tests |
:balanced | the smaller of IPOG's design and the catalog's array (Construction) where the catalog applies | the default, and covering's default engine |
:compact | that, then the row reducer (Compact) with effort | expensive tests |
Where the catalog's array has as many rows as the lower bound, no design has fewer, so IPOG isn't run. Otherwise, where the space is small (at most 100,000 combinations to cover, counted before the rules) and the catalog applies, it builds both and keeps the one with fewer rows, IPOG's on a tie, so that it gives the cases IPOG() gives unless the catalog's are fewer. Above that size it builds one: the catalog's array for a shape the catalog builds exactly (parameters that all have the same number of values, or strength + 1 parameters, with no rules, no must-include rows other than negative ones, and no stronger groups), and IPOG's design otherwise. So :balanced never returns more cases than :fast, except that above that size the catalog's array is not compared with IPOG's design: on the package's benchmarks it was never larger, which is measured, not guaranteed. :compact reduces the winner once, so it never has more rows than :balanced. The negative rows of a space with Invalid values are chosen the same way, for each invalid value.
:fast and :balanced use no randomness, and seed is not recorded. With :compact the reducer draws from a fresh generator seeded with seed (an integer of at least 0) on every call, so the same seed gives the same cases, and the result records the seed (contract §9.11). effort, a positive integer, multiplies the reducer's two budgets, of steps and of combinations read, never seconds (see Compact, which also says when a design is too large to reduce).
What Auto chose for the ordinary cases is in the result's record, cases.record.ordinary.chose, with each start it ran, its rows and what it built, in cases.record.ordinary.starts, the one kept at kept; the summary line names it, as in "Auto: Construction()". Each invalid value's choice is in cases.record.negative. The choice depends only on the request, never on the clock or a limit, but a later version may choose differently (contract §9.8). To keep a design, save it and pass it back as must_include (§9.10).
UnitTestDesign.Construction — Type
Use when every parameter has the same number of values, or there are only strength + 1 parameters, and you want the design built from a catalog of algebraic constructions: often far fewer cases than IPOG, built in milliseconds, and an orthogonal array, which shows every combination exactly once, where the catalog has one: for a prime-power number q of values, at least the strength, on at most q + 1 parameters, and for strength + 1 parameters. Other orthogonal arrays exist that it doesn't build, such as 100 cases for 4 parameters of 10 values.
Construction()The catalog engine (plan §5.4). For a space whose parameters all have the same number of values, or that has at most strength + 1 parameters, it chooses the smallest array the catalog's constructions give for that shape, from sizes alone: orthogonal arrays on prime-power numbers of values, the zero-sum array for strength + 1 parameters of any sizes, Kleitman and Spencer's arrays for two values, cover starters, products of arrays, and at strength 3 the LFSR array and its copies, small group arrays and recursions. The array is the design when the space has no rules, no must-include rows and no stronger groups. Otherwise the catalog's rows that no rule forbids seed IPOG, which adds what they leave uncovered, after the must-include rows; a catalog row that holds nothing the must-include rows and the rows before it don't is left out, so a result passed back as must_include for the same space gains no cases. Partial must-include rows are completed first, as IPOG completes them, when that leaves out more of the catalog's rows, as for a design passed back after parameters were added or removed. With stronger groups the seed is often the strongest group's own array, on that group's parameters. Above strength 3 the catalog offers only the arrays that meet the lower bound, the zero-sum and orthogonal arrays.
It refuses, with its reason, a space it has no array for, such as parameters with different numbers of values (more than strength + 1 of them): naming it is then an ArgumentError that suggests IPOG() or Auto(), which cover any request. Auto() uses it where it fits and is smaller. The result's record says which array it built, cases.record.ordinary.catalog, with its source and whether it is an orthogonal array. It uses no randomness (contract §9.4); its arrays are those of the package version, so a later version may build a smaller one (§9.8).
UnitTestDesign.Compact — Type
Use when each case is expensive to run and you want fewer of them: it removes rows from another engine's design, with the same guarantee, in a fraction of a second for most spaces.
Compact(inner; seed = 0, effort = 1)The row reducer of plan §5.3, wrapped around the covering engine inner, such as IPOG() or Construction(). It covers a request with inner, then removes rows from that design: delete the row whose removal leaves the fewest required combinations uncovered, repair the design at that row count, and repeat while each repair succeeds, stopping at the lower bound the result records (_compact). The result never has more rows than inner's and keeps its must-include rows, first and unchanged (contract §10.5). Every row it writes is checked against the rules on its own, so it never searches and its rows don't depend on feasibility_limit (§3.8). A smaller design covers fewer combinations of higher strength by accident: see the manual's Engines page. Auto(goal = :compact) runs it on the start Auto keeps: the smaller of IPOG's design and the catalog's array where it builds both, else the one it builds.
seed, an integer of at least 0, seeds a fresh generator for every call, so the same inner, seed and effort give the same rows (§9.11), and the result records the seed. effort, a positive integer, multiplies the reducer's two budgets, which count work, never seconds: at effort = 1, 30,000 repair steps or one per combination it indexes (every combination of each set of parameters that has targets, before the rules exclude any), whichever is more, and 2·10⁹ combinations read, about 10 to 20 seconds on a laptop, which strength 3 on larger spaces (30 parameters of 4 values) and strengths 4 to 6 reach. A request with more than 2^25 combinations to cover, counted before the rules, or whose start has more than 65,535 rows, gets inner's rows unreduced. The result's record says what the reducer did, cases.record.ordinary.reducer: the rows before and after, the steps, and why it stopped; and what inner did, cases.record.ordinary.start.
The negative rows of a space with Invalid values are reduced too, for each invalid value as for any request (plan §4.1). Where the inner engine refuses one of them, IPOG's rows for it are reduced.
UnitTestDesign.GND — Type
Use when you want a seeded, randomized alternative to IPOG, for instance to compare design sizes at a high strength, where it sometimes finds fewer cases; the same seed gives the same cases, and neither engine promises the smaller design.
GND(; seed = 0, candidates = 50, rng = nothing)Greedy Non-deterministic (GND). Builds each case by drawing candidates random candidate rows and keeping the one that covers the most uncovered combinations. Deterministic for a given seed (contract §9.5): each call seeds a fresh generator from seed. Pass rng to draw from a caller's generator instead; it is copied at the start of each call and never advanced (§9.6), and the recorded seed is then nothing. The 0.4 keyword M is accepted with a deprecation warning and means candidates; passing both is an error (§13.2).
candidates is a positive integer and seed an integer of at least 0 (Julia 1.10's Xoshiro refuses a negative seed), each within Int; rng is an AbstractRNG. Any other value is an ArgumentError naming the keyword, such as "candidates must be a positive integer, got 1.5".
UnitTestDesign.recommend — Function
Use when you want to see which engine Auto would use for your space, and why, before generating anything.
recommend(space; strength = 2, stronger = [], must_include = [], goal = :balanced,
feasibility_limit = 1_000_000, explanation_limit = 1_000_000) -> Recommendation
recommend(domains::NamedTuple; constraints = [], kwargs...)
recommend(name => domain, ...; constraints = [], kwargs...)
recommend(domain, domain, ...; kwargs...)What Auto(; goal) would run for this request, with one line of reason for each start it considers, the number of cases each goal can give where that is known without running, a lower bound on the cases of any design, and notes (plan §6.1). It takes the inputs and keywords that covering takes, builds the request, which checks them as covering does, and generates nothing: no combination is classified and no design is built, so it answers in about a millisecond. The limits bound the search a partial must-include row's check makes.
julia> recommend(fill(1:7, 8)...)
Recommendation: 8 parameters × 7 values, strength 2, no rules; 1372 combinations to cover
skip IPOG() not run: no design has fewer rows than the catalog's array, which meets the lower bound
use Construction() 49 rows: Bush orthogonal array, every combination exactly once
lower bound: 49 cases: the 7 × 7 = 49 combinations of p1 and p2 need a case each
goals: :fast (IPOG alone) known only after running; :balanced 49 cases, the minimum; :compact 49 cases, the minimum
covering(…) would use Construction().
julia> recommend(fill(1:6, 15)...)
Recommendation: 15 parameters × 6 values, strength 2, no rules; 3780 combinations to cover
run IPOG() covers any request
run Construction() 76 rows: Tripling (tripling from 5 columns)
rule: 3780 combinations to cover, at most 100000: both run, and the fewer rows are kept, IPOG's on a tie
lower bound: 36 cases: the 6 × 6 = 36 combinations of p1 and p2 need a case each
goals: :fast (IPOG alone) known only after running; :balanced at most 76 cases; :compact at most 76 cases
covering(…) would use the smaller of IPOG() and Construction().The goals lean toward :balanced (decision D8): it returns no more cases than :fast, which is IPOG alone, wherever it builds IPOG's design too or the catalog's array meets the lower bound, and on a space whose parameters share one number of values it is often a quarter smaller or more. Above 100,000 combinations it may build the catalog's array alone, which on the package's benchmarks was never larger than IPOG's design, though that is not guaranteed.
UnitTestDesign.Recommendation — Type
Use when you want to know, before generating, what Auto would run for your space and why, how many cases each goal can give where that is known without running, and a lower bound on the cases of any design.
RecommendationWhat recommend returns. Fields:
goal::Symbol: the goal asked about;engine::String: whatAuto(; goal)would run, such as"Construction()","IPOG()","the smaller of IPOG() and Construction()", or one of those inside"Compact(…)"for:compact, unless the reducer would return the start unreduced, as a note then says.candidates: the startsAutoconsiders, in the order a tie is broken (IPOG first), each aNamedTuple(engine, fit, runs, rows, reason): its constructor call; how it fits the request (:exact, the catalog's array is the design;:seeded, the array seeds IPOG under the rules;:native, IPOG;:unsupported); whetherAutoruns it; its number of cases when known without running (the catalog's array for an exact shape), elsenothing; and one line on why.rule::String: howAutochooses among them.lower_bound: a proven lower bound on the cases of any design for the request, andproof, why, as a result records it (TestCases'srecord);nothing, withproofsaying so, when the bound depends on the rules or the must-include rows, which only generation classifies.sizes:(fast, balanced, compact), the most cases each goal can return, where that is known without running, elsenothing.:balancedkeeps the smaller of its starts, so it is at most the catalog's array whenever it runs it, and:compactis at most:balanced.notes::Vector{String}: things about the space worth knowing before generating, such as rules that read whole cases.parameters,values(the ordinary values of each parameter),strength,stronger(asnames => strengthpairs),n_must_include,n_invalid,n_rules, andtargets, the combinations to cover before the rules exclude any.
show prints a short table:
Recommendation: 8 parameters × 32 values, strength 2, no rules; 28672 combinations to cover
skip IPOG() not run: no design has fewer rows than the catalog's array, which meets the lower bound
use Construction() 1024 rows: Bush orthogonal array, every combination exactly once
lower bound: 1024 cases: the 32 × 32 = 1024 combinations of p1 and p2 need a case each
goals: :fast (IPOG alone) known only after running; :balanced 1024 cases, the minimum; :compact 1024 cases, the minimum
covering(…) would use Construction().Analysis
Questions about a space, and measurement of any set of rows against one.
UnitTestDesign.explain — Function
Use when you want to know whether some values can appear together in a valid case, and, when they cannot, which rules exclude them.
explain(space::TestSpace, assignment; feasibility_limit = 1_000_000,
explanation_limit = 1_000_000) -> ExplanationSay whether assignment can appear in a valid row of space, and why not when it cannot (contract §1.26). assignment is a NamedTuple naming some or all parameters, or a complete Tuple or vector in parameter order. The result is an Explanation, which prints as a sentence:
julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);
julia> explain(space, (solver = :lu, tol = 1e-3))
infeasible: no valid case contains (solver = :lu, tol = 0.001); rules 1 and 2 together exclude it (rule 1: @require(mode == :exact || solver == :none); rule 2: exact mode needs a tight tolerance)
julia> explain(space, (solver = :lu,))
completable, e.g. (mode = :exact, solver = :lu, tol = 1.0e-6)An assignment with one Invalid value, at parameter p, is judged against negative rows: rules that read p do not apply (§1.27, §5.5). One with more than one is forbidden, and the result says so rather than naming a rule.
Deciding a partial assignment may search the rows that complete it, up to feasibility_limit nodes; a search that reaches the limit gives :unknown (§1.7, §3.3). When the assignment is infeasible, a deletion search looks for the rules that exclude it, within explanation_limit nodes; the assignment stays infeasible even if that search is cut short (§3.13–§3.16). The result's nodes and evaluations report the effort the answer took: nodes against those limits, and the rule checks those nodes caused, which no limit bounds.
UnitTestDesign.isallowed — Function
Use when you have one complete case and want to know whether the space's rules allow it; it evaluates the rules on that case and never searches.
isallowed(space::TestSpace, case) -> BoolWhether the complete case is a valid row of space (contract §1.25). case is a NamedTuple naming every parameter, in any order, or a Tuple or vector of values in parameter order. Values are matched by identity, and a Partition may be written by its name (§2.11).
An ordinary row is valid when no rule excludes it. A row with one Invalid value, at parameter p, is valid when no rule whose scope omits p excludes it; rules that read p are not evaluated (§5.5). A row with two or more Invalid values is never valid (§5.7).
isallowed evaluates rules on the one row and never searches (§3.9); it evaluates a lazy rule without a memo (§12.19). A partial case, an unknown name, or a value outside its parameter's domain is an ArgumentError; use explain for partial assignments.
julia> space = TestSpace((mode = [:fast, :exact], tol = [1e-3, 1e-6]);
constraints = [forbid((mode = :exact, tol = 1e-3))]);
julia> isallowed(space, (mode = :exact, tol = 1e-6))
true
julia> isallowed(space, (:exact, 1e-3))
falseUnitTestDesign.coverage — Function
Use when you have test cases from anywhere (hand-written, generated, or an older design) and want to know which combinations of values they cover and which they miss.
coverage(cases, space; strength = 2, stronger = [],
feasibility_limit = 1_000_000, explanation_limit = 1_000_000) -> Coverage
coverage(cases, domains::NamedTuple; constraints = [], kwargs...)
coverage(cases, name => domain, ...; constraints = [], kwargs...)
coverage(cases, domain, domain, ...; kwargs...)
coverage(cases::TestCases; strength, stronger, feasibility_limit,
explanation_limit) -> CoverageWhich of the combinations a set of test cases should hold it does hold (contract §1.12–§1.17): every combination of values of every strength parameters, and of every stronger group at its strength, that some valid row contains. Use it to audit a hand-written suite, to check a design built elsewhere, or to see what an edited space asks of an old result:
julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);
julia> handwritten = [(mode = :fast, solver = :none, tol = 1e-3),
(mode = :exact, solver = :lu, tol = 1e-6),
(mode = :exact, solver = :none, tol = 1e-6)];
julia> coverage(handwritten, space)
covers 8 of 11 feasible pairs, 3 missing: (mode = :exact, solver = :qr), (mode = :fast, tol = 1.0e-6), (solver = :qr, tol = 1.0e-6)
excluded: 3 pairs forbidden, 2 impossible under the constraints
julia> cases = all_pairs(space; must_include = handwritten); # keep them, add rows for the gaps
julia> coverage(cases)
covers 11 of 11 feasible pairs
excluded: 3 pairs forbidden, 2 impossible under the constraints
julia> coverage(all_triples(space; must_include = cases)) # extend the same rows to triples
covers 5 of 5 feasible triples
excluded: 7 triples forbiddencases is any collection of rows, read once: NamedTuples naming every parameter, or tuples or vectors with one value per parameter in parameter order (as a positional call such as coverage(cases, [1, 2], [:a, :b]) writes them), or a TestCases. Values match the domain by identity: 1 and 1.0 are different values (§2.1, §2.11). A row that is partial, names an unknown parameter, or holds a value outside its domain is an ArgumentError naming the row and the parameter (§1.13). The space comes in the forms covering takes; constraints = builds it from named domains.
The measurement uses the rows alone, never how they were made (§1.12):
- A row that breaks a rule, or holds more than one
Invalidvalue, is accepted, counts for nothing, and is listed inrejectedwith the rules it breaks (§1.14). Repeated rows count once (§1.11). - A combination some valid row holds is covered; no search is needed for it (§1.10). Every other combination is classified as for generation: missing (some valid row could hold it), excluded (
:forbiddendirectly by rules, or:impliedby rules together, with the rules named, §1.4), or unknown when its feasibility search reachesfeasibility_limit(§1.7). - Ordinary and negative targets are measured separately: rows with one
Invalidvalue cover only the negative targets of §6, whatever their other values, and ordinary rows only the ordinary ones (§5.9–§5.11).
The result is a Coverage, whose ordinary and negative parts hold the counts, the missing, excluded and unknown targets in target order, a breakdown per group, and the rejected rows. It prints the covered count and the missing targets. When a search ran out, it prints the known counts as bounds and lists the unresolved targets, with no percentage and no claim of completeness (§3.10); retry with a larger feasibility_limit. explanation_limit only bounds the search for which rules cause an implied exclusion (§3.13); running out leaves that attribution unresolved, never the coverage. iscomplete says whether nothing is missing or unknown.
For a TestCases, coverage(cases) measures against the result's space at its strength and stronger groups (§1.12). An excursion or a full factorial has no strength: pass one, as in coverage(cases; strength = 2). An explicit strength replaces the result's and keeps its stronger groups; a stored group whose strength is below the requested strength is an ArgumentError naming the group, since a group never asks for less than the base (§11.6). An explicit stronger replaces the stored groups, and stronger = [] drops them.
Coverage describes the rows given. To measure the cases that ran or passed, pass those rows (§1.17).
See also missing_interactions, iscomplete.
UnitTestDesign.Coverage — Type
Use when you read what coverage measured: the covered, missing, excluded and unresolved combinations, with ordinary and negative targets kept apart.
CoverageWhat coverage measured: which of the requested combinations the supplied rows contain (contract §1.12–§1.17). Fields:
ordinary,negative: aCoverageParteach, measured separately (§5.10). The negative part is empty unless the space hasInvalidvalues.space: theTestSpacemeasured against.strength,stronger: the request,strongerasnames => strengthpairs without the base group, as inTestCases.limits:(feasibility_limit = …, explanation_limit = …), the budgets the classification used (§3.3, §3.13).
iscomplete(c) is true when no target of either part is missing or unknown (§1.16). An unknown target makes it false without showing that anything is missing: the counts are then bounds, and nothing prints a percentage (§3.10).
It prints as a sentence per part (the negative part only when the space has Invalid values), then the excluded counts, duplicates and rejected rows. For the space and rows of coverage's example, with the second row given twice and a row that breaks a rule third:
julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);
julia> handwritten = [(mode = :fast, solver = :none, tol = 1e-3),
(mode = :exact, solver = :lu, tol = 1e-6),
(mode = :exact, solver = :none, tol = 1e-6)];
julia> rows = [handwritten[1], handwritten[2], (mode = :fast, solver = :lu, tol = 1e-3),
handwritten[3], handwritten[2]];
julia> coverage(rows, space)
covers 8 of 11 feasible pairs, 3 missing: (mode = :exact, solver = :qr), (mode = :fast, tol = 1.0e-6), (solver = :qr, tol = 1.0e-6)
excluded: 3 pairs forbidden, 2 impossible under the constraints
1 duplicate row counted once
1 row rejected: row 3 breaks rule 1 (@require(mode == :exact || solver == :none))and, when a search reached its limit, gives bounds and the unresolved targets instead of an exact count:
covers 0 of at least 0 feasible pairs; 448 pairs unresolved (feasibility_limit = 1): (x1 = 1, x2 = 1), …, and 438 more; no exact percentageUnitTestDesign.iscomplete — Function
Use when a test should assert that a set of cases covers every feasible combination, as in @test iscomplete(coverage(cases, space)); it is false when anything is missing or unresolved.
iscomplete(c::Coverage) -> Booltrue when every feasible target, ordinary and negative, is covered and no target is unresolved (contract §1.16). An unknown target makes it false: coverage is never claimed complete under an exhausted limit (§1.7, §3.10). Rejected rows do not change it; they are listed in the result (§1.14).
A rejected row does not make a result incomplete. To test that committed cases are still valid, check the rows with isallowed as well, @test all(case -> isallowed(space, case), cases); otherwise a new rule that forbids a committed row passes unnoticed.
UnitTestDesign.missing_interactions — Function
Use when you want the list of feasible combinations your cases miss, to add cases for them; it throws rather than return a list that a search limit left uncertain.
missing_interactions(cases, space; strength = 2, stronger = [],
feasibility_limit = 1_000_000, explanation_limit = 1_000_000)
missing_interactions(cases::TestCases; kwargs...)The feasible combinations that no valid row of cases holds, in target order: coverage(cases, space; ...)'s missing targets, ordinary then negative (contract §3.11). It takes the arguments of coverage.
An empty list means nothing is missing: it is returned only when every target was resolved. If a feasibility search reached feasibility_limit, some combination could be feasible and missing without anyone knowing, so missing_interactions throws a ResourceLimitError instead; call coverage with the same arguments for the missing targets known so far and the unresolved ones, or raise the limit.
With the space and rows of coverage's example:
julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);
julia> handwritten = [(mode = :fast, solver = :none, tol = 1e-3),
(mode = :exact, solver = :lu, tol = 1e-6),
(mode = :exact, solver = :none, tol = 1e-6)];
julia> missing_interactions(handwritten, space)
3-element Vector{NamedTuple}:
(mode = :exact, solver = :qr)
(mode = :fast, tol = 1.0e-6)
(solver = :qr, tol = 1.0e-6)UnitTestDesign.report — Function
Use when you want to check what a generated result promises and see the evidence: the guarantee, with its coverage figures measured from the rows, what the rules excluded and why, bonus coverage at the next strength, and how coverage grows over the first cases.
report(cases::TestCases; feasibility_limit = 1_000_000,
explanation_limit = 1_000_000) -> ReportCheck what a generated result promises and say it in one line, with the evidence (contract §1.23): the guarantee, its coverage figures measured from the rows; the targets excluded, each with the rules that exclude it; bonus coverage at the next strength; the prefix curve, how much the first rows cover, for suites that run only part of the cases; and the seed. show prints only what generation recorded; report recounts.
julia> space = TestSpace(
(mode = [:fast, :exact], solver = [:none, :lu, :qr], tol = [1e-3, 1e-6]);
constraints = [
@require(mode == :exact || solver == :none),
forbid((mode = :exact, tol = 1e-3); reason = "exact mode needs a tight tolerance"),
]);
julia> report(all_pairs(space))
5 cases cover all 11 feasible pairs of a 12-combination space (3 pairs forbidden, 2 impossible under the constraints)
excluded:
(mode = :fast, solver = :lu): forbidden by rule 1 (@require(mode == :exact || solver == :none))
(mode = :fast, solver = :qr): forbidden by rule 1 (@require(mode == :exact || solver == :none))
(mode = :exact, tol = 0.001): forbidden by rule 2 (exact mode needs a tight tolerance)
(solver = :lu, tol = 0.001): impossible because rules 1 and 2 combine (rule 1: @require(mode == :exact || solver == :none); rule 2: exact mode needs a tight tolerance)
(solver = :qr, tol = 0.001): impossible because rules 1 and 2 combine (rule 1: @require(mode == :exact || solver == :none); rule 2: exact mode needs a tight tolerance)
size: 5 cases; lower bound 4: the 4 feasible combinations of mode and solver need a case each
bonus: 5 of 5 feasible triples covered
prefix curve:
first 1 of 5 cover 27% (3 of 11)
first 2 of 5 cover 54% (6 of 11)
first 3 of 5 cover 72% (8 of 11)
first 4 of 5 cover 90% (10 of 11)
first 5 of 5 cover 100% (11 of 11)
seed: none (Auto uses no randomness)A covering result is measured at its strength and stronger groups. An excursion or a full factorial has no strength, so it is measured at strength min(2, number of parameters) and the guarantee says so; an excursion is not a covering design, and the guarantee says that too (§1.12, §7.7). An excursion's must-include rows are kept first and are not bound by its distance, so the guarantee counts them apart, as in "1 must-include row kept first, then 1 case within distance 0 of (…)" (§7.5, §7.9).
With Invalid values the rows make two guarantees, stated apart (§5.10): the ordinary one, then, after "negative:", what the negative rows cover of the negative targets (§6) and which of those are excluded, as in "6 cases cover all 9 feasible pairs of a 12-combination space (2 pairs forbidden, 1 impossible under the constraints); negative: covers 4 of 4 feasible pairs". The excluded list gives the negative exclusions after the ordinary ones. The bonus line and each prefix-curve line give the negative figure beside the ordinary one, which is labeled ordinary, so a prefix that covers every ordinary pair does not look complete while negative targets remain:
prefix curve:
first 1 of 3 cover 100% of ordinary pairs (1 of 1); negative 0 of 2
first 2 of 3 cover 100% of ordinary pairs (1 of 1); negative 1 of 2
first 3 of 3 cover 100% of ordinary pairs (1 of 1); negative 2 of 2The measurement searches, within feasibility_limit nodes per target, for the targets no row holds (§3.9). A search that runs out leaves its target unresolved: the counts become bounds ("8 of at least 11"), no percentage is printed, and nothing is called complete (§3.10, §3.12). explanation_limit bounds the search for the rules behind an implied exclusion (§3.13); running out marks that explanation unresolved and changes no count (§3.15).
The report is the verification (§1.23). The exclusions it lists, with their rules and whether each explanation is verified inclusion-minimal, are its own, found with its own feasibility_limit and explanation_limit, not copied from generation: a result generated with explanation_limit = 1 and reported with the default shows verified explanations, and one reported with explanation_limit = 1 shows the explanations that limit left unresolved, whatever generation found. Exclusions recorded at generation are a fallback, used only for targets this measurement left unknown, and each is printed with "(recorded at generation)"; they are in the recorded field, apart from excluded.
UnitTestDesign.Report — Type
Use when you read the fields of what report found: the guarantee, the coverage, the exclusions, bonus coverage, and the prefix curve.
ReportWhat report found about a TestCases. Fields:
guarantee::String: the claim the rows meet, its coverage figures measured from the rows, such as "5 cases cover all 11 feasible pairs of a 12-combination space (3 pairs forbidden, 2 impossible under the constraints)". The rest of the line is what the result recorded at generation (§1.19): the must-include rows kept first, an excursion's distance, base, dropped rows and values that never appear, a full factorial's "every valid row", and a randomized engine's seed, as its record's configuration names the engine.strategy::Symbol,n_cases::Int,engine::Symbol,seed,n_must_include::Int,record::NamedTuple: as the result recorded them (TestCases);recordholds a covering design's lower bound, with its proof and whether the rows meet it, whichshowprints on its "size:" line, the engine's configuration, from which the guarantee and the seed line name the engine and the call that repeats the cases, and the stages that ran.strength::Int: the strength measured: the result's, ormin(2, number of parameters)for an excursion or a full factorial, which have none (contract §1.12).coverage::Coverage: the verification, atstrengthand the result'sstrongergroups.excluded::Vector{Exclusion}: the targets no valid row can hold, with the rules that exclude them, as this report's measurement classified and explained them with its ownexplanation_limit:coverage.ordinary.excluded, thencoverage.negative.excluded, whose targets hold anInvalidvalue (§1.23, §5.10).recorded::Vector{Exclusion}: exclusions generation recorded for targets this measurement left unresolved (a search reachedfeasibility_limit), ordinary then negative, shown as "(recorded at generation)". Empty when the measurement resolved every target, and always for an excursion or a full factorial, which record none.bonus: coverage of the same rows atstrength + 1, as(strength, covered, feasible, unknown, negative, applicable, reason).covered,feasibleandunknowncount the ordinary targets;negativeis(covered, feasible, unknown)for the negative targets atstrength + 1(§6), all 0 when the space has noInvalidvalues.applicableisfalse, with areason, when there are no targets atstrength + 1: the strength already equals the number of parameters (§3.12).prefix: the prefix curve, one(cases, covered, feasible, unknown)per prefix length1:n_cases: how many ordinary targets atstrengththe firstcasesrows cover.prefix_negative: the same for the negative targets, from the same pass over the rows: how many of them the firstcasesrows cover. Itsfeasibleis 0 throughout when the space has noInvalidvalues.
Every progress figure keeps the two parts apart (§5.9, §5.10): ordinary rows cover only ordinary targets and negative rows only negative ones, so the ordinary prefix curve can reach 100% while the negative targets are not yet covered. With Invalid values, show labels the ordinary figures as ordinary and prints the negative ones beside them, as in "first 1 of 3 cover 100% of ordinary pairs (1 of 1); negative 0 of 2" and "bonus: 5 of 6 feasible triples covered; negative: 3 of 4".
With unknown targets (a search reached feasibility_limit), feasible in coverage, bonus and the prefix curves is a lower bound, and nothing prints a percentage or claims completeness (§3.10, §3.12).
Report fields are plain data except coverage.space, the TestSpace, which holds the rules' predicates. UnitTestDesign.plain(report) gives a representation with no executable state: nested NamedTuples and Vectors of Int, Float64, String, Symbol, Bool and nothing, with the space reduced to its names, its printed domains and its rule labels. github_matrix validates values itself and rejects wrappers rather than stringifying them.
UnitTestDesign.design_sizes — Function
Use when choosing a strategy before committing to one: it shows how many cases each strategy produces for your space, and what each covers.
design_sizes(space; strengths = 1:3, distances = 1:2, engine = Auto(), limit = 10^6,
from = nothing, feasibility_limit = 1_000_000,
explanation_limit = 1_000_000) -> DesignSizes
design_sizes(space; engine = [IPOG(), Auto(goal = :compact)], kwargs...)
design_sizes(domains::NamedTuple; constraints = [], kwargs...)
design_sizes(name => domain, ...; constraints = [], kwargs...)
design_sizes(domain, domain, ...; kwargs...)How many cases each strategy gives for space, before committing to one (plan Phase 5 step 4). Runs the full factorial, covering at each of strengths up to the number of parameters (larger ones are left out), and excursions at each of distances from from (the first value of each parameter when omitted), and measures each design's pairs and triples with coverage:
julia> design_sizes(:n => [1, 2, 3], :level => ["low", "mid", "high"],
:tol => [1.0, 3.7, 4.9], :kind => [:greedy, :relax, :optim])
strategy cases share pairs triples
full_factorial 81 100.0% 54/54 108/108 valid 81 of 81
covering(1) 3 3.7% 18/54 12/108
covering(2) 9 11.1% 54/54 36/108
covering(3) 27 33.3% 54/54 108/108
excursions(1) 9 11.1% 30/54 28/108
excursions(2) 33 40.7% 54/54 76/108
case counts are the rows each strategy produced with Auto, not lower boundsshare is the fraction of the valid rows, known when the full product is at most limit and the full factorial counted them; above it the table gives the product only and no shares (§7.3). A strategy that stops at a resource limit is shown with its status in place of a count (§3.12); the others still run. The counts are the rows each strategy produced with engine, not lower bounds (§8.3). The space comes in the forms covering takes.
engine may be a vector of engines, to compare them: each covering strength then has one row per engine, in the order given, under an engine column. An engine that does not cover a request, such as Construction on parameters with different numbers of values, shows its reason in place of a count, as a resource limit does.
With Invalid values each count is two figures, ordinary + negative (§5.10): cases as "4 + 3", the ordinary rows and the negative rows, and pairs and triples as "9/9 + 4/4", the ordinary targets covered of the feasible ones and then the negative targets (§6). A line under the table says so. The share is of all valid rows, ordinary and negative.
UnitTestDesign.DesignSizes — Type
Use when you read the table design_sizes returns: one row per strategy, with its case count, its share of the valid cases, and its coverage.
DesignSizesThe table design_sizes returns. Fields: parameters (the names), total (the full product), valid (the valid rows, ordinary and negative, or nothing when the product is above limit and they were not counted), engine (the name the results record of the first engine, such as :Auto), limit, has_invalid (whether the space has Invalid values, so that each figure has a negative part), rows, one per strategy run, and engines, every engine compared, as its constructor call ("IPOG()", "Auto(goal = :compact)"). Each row is a NamedTuple:
strategy:"full_factorial","covering(s)"or"excursions(d)";kind(:full_factorial,:covering,:excursion) andlevel(the strength or distance, 0 for the full factorial).status::ok;:resource_limitwhen the strategy stopped at a limit, withmessagenaming it, and no case count or share (§3.12);:unsupportedwhen the engine does not cover the request, withmessagegiving its reason, asConstruction()refuses mixed value counts; or:invalid_basefor an excursion whose default base breaks a rule.cases: the rows the strategy produced, ordinary and negative, andshare,cases / valid(nothingwhenvalidis unknown).pairs,triples: the design's coverage of the ordinary targets at strength 2 and 3 as(covered, feasible, unknown);nothingwhen the space has too few parameters or the strategy did not finish.negative_cases: how many ofcasesare negative rows, with oneInvalidvalue;negative_pairs,negative_triples: the coverage of the negative targets (§6) at strength 2 and 3, in the same form. 0 and zero counts when the space has noInvalidvalues;nothingwherecasesorpairsandtriplesare.engine: for a covering row, the engine that made it, as inengines, the call the result records (record.engine.callof itsTestCases);nothingfor the full factorial and the excursions, which use none.
Every figure keeps the ordinary and negative parts apart (§5.9, §5.10). Case counts are what the engine produced, not lower bounds (§8.3).
Types returned by explain and coverage
These types are not exported; their fields are part of the results above.
UnitTestDesign.Explanation — Type
ExplanationThe result of explain: why an assignment is or is not part of a valid row (contract §1.26). It prints as one sentence. Fields:
assignment: the assignment explained, as aNamedTuplein parameter order, with values as the domain stores them.outcome: one of:allowed, a complete, valid row;:forbidden, excluded directly:rulesnames every rule whose scope is entirely assigned and which excludes the assignment. An assignment with more than oneInvalidvalue is forbidden with no rule (§1.27);:completable, part of the valid rowwitness;:infeasible, part of no valid row, though no rule excludes it directly:rulesis a set of rules proven to exclude it together (§1.4);:unknown, undecided withinfeasibility_limit(§1.7).
rules: positions of the rules in the space'sconstraints, in order.labels: each rule's label: its reason, its macro source text, or its position and scope (§12.3).minimal: for:infeasible,:verifiedwhen removing any one ofruleswas shown to make the assignment completable, and:unresolvedwhen a limit stopped that check; thenrulesis still sufficient, but may hold a rule it does not need (§3.15, §3.16).:not_applicableotherwise.witness: for:completableand:allowed, a valid row containing the assignment, as aNamedTuple; otherwisenothing.limit: the limit that decided an:unknownoutcome or an:unresolvedexplanation, askeyword => value; otherwisenothing.nodes,evaluations: the search effort of this answer.nodescounts the tentative assignments of §3.3, those of the feasibility search plus those of the deletion search;feasibility_limitandexplanation_limitbound them.evaluationscounts rule checks, each one consultation of one rule on one assignment of its scope (a table lookup, or a memoized evaluation for a lazily evaluated rule): the direct check, forward checking, and the deletion trials. No limit boundsevaluations; it shows how much rule checking the nodes caused (§3.3).
UnitTestDesign.CoveragePart — Type
CoveragePartThe ordinary or the negative half of a Coverage (contract §1.15, §5.10). Ordinary targets are combinations of ordinary values, covered only by valid ordinary rows; negative targets hold one Invalid value and are covered only by valid negative rows (§1.9, §5.9, §6). Fields:
covered::Int: targets some valid row of this kind contains (§1.9).feasible::Int:coveredplus the missing targets. Whenunknownis empty this is the number of feasible targets; otherwise it is a lower bound, since an unresolved target may be feasible too (§3.10).missing::Vector{NamedTuple}: feasible targets no valid row contains, in target order (§9.7). Each is a partial row in parameter order.excluded::Vector{Exclusion}: targets no valid row can contain,:forbiddenor:implied, with the rules that exclude them, in target order (§1.4).unknown::Vector{NamedTuple}: targets whose feasibility search reachedfeasibility_limit: neither missing nor excluded (§1.7, §3.10).groups: one entry per requested group, the base group first, then eachstrongergroup in the order the request keeps them:names,strength, and the group's owncovered,feasible,missing_count,excluded_countandunknown_count. A target that two groups share counts in each group's entry and once in the totals above (§1.8).rows::Int: distinct valid rows of this kind;duplicates::Int: valid rows that repeat an earlier one and count once (§1.11).rejected: rows accepted as input that contribute nothing (§1.14), as(index, row, reason, rules):indexis the row's position among the cases,rowthe row as given,reason:violates_rule(withrules, every applicable rule it breaks, in rule order) or:multiple_invalid(more than oneInvalidvalue;rulesempty). A row with anInvalidvalue is rejected in the negative part, any other in the ordinary part.
After the run
diagnose and followups are experimental: their results are hypotheses to test, not proof. github_matrix writes cases for a CI workflow.
UnitTestDesign.diagnose — Function
Use when some cases failed and you want ranked hypotheses about which values, or combinations of values, the failures have in common. Experimental: the ranking is a set of hypotheses, not proof.
diagnose(cases, passed::AbstractVector{Bool}; strength = nothing, space = nothing) -> DiagnosisThe ranking is a set of hypotheses, not proof (contract §8.6). Its interface may change in a minor release.
passed[k] is the outcome of cases[k]: true for a pass, false for a failure. Collect outcomes however you like (a Test.jl loop, a cluster job, a spreadsheet); diagnosis is a pure function of the cases and outcomes, and runs nothing (§14.1).
A suspect is a combination of 1 to strength values that appears in at least one failing case and in no passing case. Suspects are ranked by the number of failing cases that contain them, most first; then smaller combinations first, since a single value that explains the failures is a simpler hypothesis than a pair; then by parameter order and domain order, so the ranking is deterministic. Suspects with the same failure pattern, those that occur in exactly the same failing cases, are grouped, and the display lists them as "same failures as" the group's first member: these outcomes cannot tell them apart, and only new cases can. followups proposes those cases.
julia> space = TestSpace((n = [10, 100, 1000, 10000], method = [:newton, :bicg, :gmres],
tol = [1e-3, 1e-6], sparse = [false, true]));
julia> cases = all_pairs(space);
julia> passed = [!(c.method == :newton && c.sparse) for c in cases]; # the bug
julia> diagnose(cases, passed)
3 failures of 12 cases; 6 suspects in 5 groups (hypotheses, not proof)
1. (method = :newton, sparse = true) — in 3 of 3 failures
2. (method = :newton, tol = 0.001) — in 2 of 3 failures
3. (n = 10, method = :newton) — in 1 of 3 failures
4. (n = 100, method = :newton) — in 1 of 3 failures
5. (n = 10000, method = :newton) — in 1 of 3 failures — same failures as (n = 10000, sparse = true)The true cause ranks first, in every failure. The other suspects appeared only in failing cases too, so nothing yet says whether they work; followups finds a case for each that holds it and no other suspect.
Read the ranking as hypotheses (§8.6):
- Several faults at once split the failures between their causes, so each cause is in fewer failures than all of them, and a combination that happens to share failing cases with two faults can rank above both.
- An intermittent failure makes a passing case look innocent when it is not; the true cause may then be missing from the list, since a suspect must appear in no passing case.
- A fault that needs more values together than
strengthleaves no suspect, or only combinations that happen to occur with it. Diagnose again at a higher strength. - A wrong answer and a thrown error are both
false; diagnose them separately if they may have different causes.
Every suspect is unverified by passing cases by definition: each one's passes is 0.
cases is a TestCases, or any other collection of rows, such as a vector, read once. For a TestCases, the space is the result's, and strength defaults to the result's strength, or to min(2, number of parameters) for an excursion or a full factorial, which have none. For other rows, pass the space they belong to as space (a TestSpace); strength then defaults to min(2, number of parameters). Rows are NamedTuples naming every parameter, or tuples or vectors in parameter order, and values match the domain by identity (§2.1). Give the labeled rows, before realize: a drawn value is not in the domain. A positional result's parameters are named p1, p2, ….
diagnose takes the cases and outcomes as observed (§8.6). A case that breaks a rule, or holds more than one Invalid value, is ranked like any other. A failing case that broke the rules can leave a suspect that no valid case holds, which followups reports as :inseparable with no other suspects.
length(passed) must equal the number of cases. Each row is counted once per appearance, so a case that ran twice and failed twice is in two failures. The result's status says how the comparison went:
:no_failures: nothing failed, so there is nothing to diagnose.:all_failed: every case failed, so no combination is implicated over another; check the setup before suspecting the cases.:conflicting: some case appears more than once with different outcomes. Those rows are listed inconflictingand left out of the comparison, which proceeds on the rest.:ranked: failing and passing cases were compared.
Outcomes enter the package here and nowhere else: coverage describes the rows it is given, whatever their outcomes, so to see what the passing cases covered, pass those rows to it (§1.17).
UnitTestDesign.followups — Function
Use when diagnose has ranked suspects and you want the next cases to run: for each suspect it searches for a valid case that holds that suspect and no other, and says when no such case exists or the search ran out. Experimental.
followups(d::Diagnosis; feasibility_limit = 1_000_000, explanation_limit = 1_000_000,
prefer = :nearest) -> Vector{Followup}Follow-up cases test hypotheses; they carry no minimum-distance guarantee (contract §8.6). The interface may change in a minor release.
For each suspect, followups searches for a valid case that holds it and no other suspect, so that the case's outcome speaks to that suspect alone. Such a case need not exist, and a search may reach its limit first, so isolation is not guaranteed: the status says which. The result has one Followup per suspect, in the diagnosis's rank order, and each has a status:
:found:caseis such a case, aNamedTuplenaming every parameter (for a positional result,Tuple(f.case)is the positional form), andkindsays whether it is:ordinaryor:negative. A negative case holds anInvalidvalue, so run it as the negative test it is.:indistinguishable: the suspect contains another suspect, listed inothers, so every case that holds it holds that one too. No case can isolate it, whatever the rules; this is a matter of construction, not a search. When the two also fail in the same cases, as a value and a pair that contains it can, the outcomes so far cannot tell them apart either. The smaller suspect's follow-up is the case to run.:inseparable: proven by exhausted searches: under the space's rules, no valid case of any kind, ordinary or negative, holds the suspect without another suspect.searchedlists the kinds of case proven, andproofsholds each kind's proof, aFollowupProof: the other suspects and the space's rules that suffice to exclude an isolating case of that kind, asexplainnames them (§3.13–§3.16), and whether each part was verified necessary for that kind.others,rulesandlabelsare the union of the proofs, sufficient for every kind but not necessarily minimal as a whole, since each kind's search runs on its own. Sominimalis judged per kind of case::verifiedsays that every kind's proof is minimal for that kind, not that the union is. When the kinds' proofs differ, the printed line gives each one instead of the union. With noothers, no valid case holds the suspect at all: the failing cases broke the rules.:unknown: the search reachedfeasibility_limitbefore deciding (§3.17). Retry with a larger limit.
julia> space = TestSpace((n = [10, 100, 1000, 10000], method = [:newton, :bicg, :gmres],
tol = [1e-3, 1e-6], sparse = [false, true]));
julia> cases = all_pairs(space);
julia> d = diagnose(cases, [!(c.method == :newton && c.sparse) for c in cases]);
julia> followups(d)
6-element Vector{UnitTestDesign.Followup}:
(method = :newton, sparse = true): found (n = 1000, method = :newton, tol = 1.0e-6, sparse = true), 1 change from case 12
(method = :newton, tol = 0.001): found (n = 1000, method = :newton, tol = 0.001, sparse = false), 2 changes from case 9
(n = 10, method = :newton): found (n = 10, method = :newton, tol = 1.0e-6, sparse = false), 1 change from case 12
(n = 100, method = :newton): found (n = 100, method = :newton, tol = 1.0e-6, sparse = false), 2 changes from case 11
(n = 10000, method = :newton): found (n = 10000, method = :newton, tol = 1.0e-6, sparse = false), 2 changes from case 9
(n = 10000, sparse = true): found (n = 10000, method = :bicg, tol = 0.001, sparse = true), 1 change from case 9Each search is the witness search of explain, over the space's rules plus one temporary forbidden combination per other suspect. Those isolation conditions apply to every case, including one whose invalid value is at a parameter they name (§3.17). A case is valid under its own kind's rules (§5.3–§5.5), so every kind that could hold the suspect is searched:
- A suspect with no
Invalidvalue: ordinary cases first, then, for each parameter the suspect leaves out and each invalid value of that parameter, in parameter and domain order, negative cases with that value, where the rules that read its parameter do not apply (§5.5). A rule that keeps the suspect out of every ordinary case does not make it inseparable when a negative case holds it alone. - A suspect with one
Invalidvalue: negative cases with that value. - A suspect with two is
:inseparable, since no valid case holds two (§5.7).
With prefer = :nearest, each search starts from a failing case that holds the suspect: each parameter tries that case's value first, so the case found tends to change few of its values. It starts in turn from each of the first five such failing cases, for each kind of case, and keeps the case with the fewest changes; on a tie, the kind searched first, so an ordinary case before a negative one. This is a heuristic with no minimum-distance guarantee (§8.6). With prefer = :domain, values are tried in domain order, and the first kind with a case gives it, ordinary before negative. Either way from is the failing case that the case found is closest to, and changes the number of parameters that differ from it.
feasibility_limit bounds each search (§3.3); explanation_limit bounds the deletion search that names an :inseparable suspect's reasons, and running out leaves the reasons :unresolved, never the status (§3.15).
UnitTestDesign.github_matrix — Function
Use when a GitHub Actions workflow should run one job per test case: write the cases as the JSON document {"include": [...]}, one object per case, for a workflow to read with fromJSON.
github_matrix(cases; io = stdout)
github_matrix(io::IO, cases)cases is a TestCases or a vector of rows. Each object's keys are the parameter names; a positional result's parameters are p1, p2, …, and so are a Tuple row's. A NamedTuple row supplies its own names. The document is one line with no trailing newline; sprint(github_matrix, cases) returns it as a String. The function returns nothing.
Each name must be valid UTF-8 and nonempty, or the call is an ArgumentError naming the row and the field name. JSON.jl escapes quotes, backslashes and control characters in names as in values. A job reads a value as matrix.<name> when the name starts with a letter or _ and holds only letters, digits, - and _, as most Julia names do; any other name, such as one with a space, a leading digit or a ., still works through index syntax, matrix['<name>'], so github_matrix does not reject it.
Values are written as JSON can hold them:
| Value | JSON |
|---|---|
String (any valid UTF-8 AbstractString) | string |
Symbol | string, its name: :qr becomes "qr" |
Bool | true or false |
finite Integer or AbstractFloat | number: 1.0e-6 reads back as a Float64 |
nothing | null |
Anything else is an ArgumentError naming the row and the field: NaN and Inf, missing, an Invalid or Partition wrapper, a vector or other collection, a Char, and any other type. Map such values to supported ones first, for example with [merge(row, (tol = string(row.tol),)) for row in cases]. Every row is checked before anything is written, so after an error io holds nothing from this call.
GitHub Actions runs at most 256 jobs from one matrix; above that, github_matrix still writes every row and warns once.
A workflow computes the matrix in one job and fans out in the next:
jobs:
design:
runs-on: ubuntu-latest
outputs:
cases: ${{ steps.gen.outputs.cases }}
steps:
- uses: julia-actions/setup-julia@v2
- run: julia -e 'using Pkg; Pkg.add("UnitTestDesign")'
- id: gen
run: |
echo "cases=$(julia -e 'using UnitTestDesign; github_matrix(all_pairs((os = ["ubuntu-latest", "macos-latest"], solver = [:lu, :qr], threads = [1, 4])))')" >> "$GITHUB_OUTPUT"
test:
needs: design
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.design.outputs.cases) }}
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v4
- uses: julia-actions/setup-julia@v2
- run: julia --project -e 'using Pkg; Pkg.test()'
env:
SOLVER: ${{ matrix.solver }}
JULIA_NUM_THREADS: ${{ matrix.threads }}fromJSON turns the document into a matrix whose include list is the cases, and each job reads its values as matrix.<parameter>. GitHub formats numbers itself when it substitutes them; write a value as a string when the job needs its exact text.
julia> cases = all_pairs((os = ["ubuntu-latest", "macos-latest"], solver = [:lu, :qr], threads = [1, 4]));
julia> github_matrix(cases)
{"include":[{"os":"ubuntu-latest","solver":"lu","threads":1},{"os":"macos-latest","solver":"qr","threads":1},{"os":"macos-latest","solver":"lu","threads":4},{"os":"ubuntu-latest","solver":"qr","threads":4}]}Types returned by diagnose and followups
These types are not exported; their fields are part of the results above.
UnitTestDesign.Diagnosis — Type
DiagnosisThe result of diagnose. Fields:
status::Symbol::ranked(failing and passing cases were compared; the list may be empty),:no_failures(nothing failed; no suspects),:all_failed(every case failed, so no combination is implicated over another; no suspects), or:conflicting(some case appears more than once with different outcomes; those rows are listed inconflictingand left out, and the rest are compared as for:ranked).groups::Vector{Vector{Suspect}}: the suspects, grouped by failure pattern, the set of failing cases that contain them. The members of a group have the same failure pattern, so these outcomes cannot tell them apart;showlists them as "same failures as" the first. This is not the:indistinguishablestatus offollowups, which means that one suspect contains another. Groups are ranked by the number of failing cases, most first, then by their representative, the group's first and smallest member: fewer values first, then parameter order, then domain order. Members follow the same order.suspects::Vector{Suspect}: every suspect, group by group, in rank order.strength::Int: the largest combination considered.n_cases::Int,n_failed::Int: the rows given, and how many failed, conflicting rows included.failing::Vector{Int},passing::Vector{Int}: the positions of the failing and passing rows compared, conflicting rows left out.conflicting::Vector{Vector{Int}}: the positions of rows that hold the same case with different outcomes, one vector per case.space::TestSpace: the space the rows belong to, forfollowups.rows::Vector{Vector{Int}}: every row in index space (internal).
UnitTestDesign.Suspect — Type
SuspectA combination diagnose implicates: it appears in at least one failing case and in no passing case. A hypothesis, not a finding (§8.6).
combination::NamedTuple: its values, in parameter order, as the domain stores them (wrappers kept).failures::Int: the number of failing cases that contain it.passes::Int: the number of passing cases that contain it. Always 0: a combination seen in a passing case is not a suspect. It is kept so that every count reads as what it is, "in 3 failures, 0 passes", and no reader mistakes a suspect for a combination known to fail.failing::Vector{Int}: the positions of those failing cases among the rows given, ascending.key::Vector{Int}: the combination in index space, one value index per parameter and 0 for a parameter it leaves out (internal).
UnitTestDesign.Followup — Type
FollowupOne suspect's follow-up from followups. Fields:
suspect::NamedTuple: the suspect's combination.status::Symbol::found,:indistinguishable,:inseparable, or:unknown(seefollowups).case: for:found, a valid case holding the suspect and no other suspect, as aNamedTuplenaming every parameter; otherwisenothing.kind::Symbol: for:found, the kind ofcase::ordinary, or:negativewhen it holds oneInvalidvalue and is valid under the negative-row policy (§5.3, §5.5).:noneotherwise.from::Int,changes::Int: for:found, the failing case (its position among the rows diagnosed) thatcaseis closest to, and how many parameters differ from it; 0 changes means that failing case already holds this suspect and no other. 0 and -1 otherwise.others::Vector{NamedTuple}: for:indistinguishable, the suspects this one contains; for:inseparable, the other suspects in any kind's proof, in rank order, of which every valid case holding this one holds at least one. Otherwise empty.rules::Vector{Int},labels::Vector{String}: for:inseparable, the space's rules in any kind's proof, as positions in itsconstraints, ascending, and their labels.minimal::Symbol: for:inseparable, combined fromproofs::unresolvedwhen any proof is,:verifiedwhen every proof is, and:not_applicableotherwise, so a kind with a direct proof (:not_applicable) makes it:not_applicableunless another kind is unresolved. Minimality is judged within each kind of case::verifiedsays that each rule and suspect was shown to be needed for the kind whose proof it is part of, not that the union is minimal.:not_applicablewhen there are no proofs, and for every other status.limit:keyword => valuefor the limit that made the status:unknown, or, for:inseparable, that left a proof:unresolved(the first such proof, in the order ofsearched); otherwisenothing.searched::Vector{NamedTuple}: the kinds of case searched, in the order searched:NamedTuple()for ordinary cases, and(p = v,)for negative cases with the invalid valuevatp. For:inseparable, every kind that could hold the suspect, each proven to hold no isolating case. Empty for:indistinguishableand for a suspect with twoInvalidvalues.proofs::Vector{FollowupProof}: for:inseparable, oneFollowupProofper kind searched, in the same order, soproofs[i].searched == searched[i].others,rulesandlabelsabove are the union of the proofs: sufficient for every kind, but not necessarily minimal as a whole, since each kind's deletion search runs on its own. Empty for every other status, and for a suspect with twoInvalidvalues, which is:inseparablewithout a search. A:foundsuspect has none even when a kind searched before the case was found was proven.
UnitTestDesign.FollowupProof — Type
FollowupProofOne kind of case's proof in an :inseparable Followup: no valid case of that kind holds the suspect without another suspect (§3.17). Fields:
searched::NamedTuple: the kind of case, asFollowup.searchednames it:NamedTuple()for ordinary cases, and(p = v,)for negative cases with the invalid valuevatp.rules::Vector{Int},labels::Vector{String}: the space's rules in the proof, as positions in itsconstraints, ascending, and their labels. Only rules that apply to this kind of case appear (§5.5, §5.6).others::Vector{NamedTuple}: the other suspects in the proof, in rank order, of which every valid case of this kind holding the suspect holds at least one.minimal::Symbol::verifiedwhen each rule and suspect in the proof was shown to be needed for this kind of case (§3.16),:unresolvedwhen a limit stopped that check (the proof still stands, §3.15), and:not_applicablewhen the proof is direct: the suspect's values, with this kind's invalid value, already break each rule listed and hold each suspect listed, so no deletion search ran.limit: for:unresolved, the limit that stopped the check, askeyword => value; otherwisenothing.
Deprecated
Each warns through Base.depwarn and will be removed in the next breaking release. The deprecated keywords (n_way, seeds, wayness, and GND(M = ...)) are described with the functions that accept them.
UnitTestDesign.all_tuples — Function
Deprecated alias of covering: call covering with the same inputs and keywords, writing strength for n_way.
all_tuples(input...; n_way = 2, kwargs...)It warns through Base.depwarn and will be removed in the next breaking release (contract §13.1, §13.2). See the migration table in the manual.
UnitTestDesign.values_excursion — Function
Deprecated alias of excursions at distance 1: call excursions(input...; distance = 1).
values_excursion(input...; kwargs...)It takes the keywords of excursions, and the 0.4 n_way as the distance. It warns through Base.depwarn and will be removed in the next breaking release (contract §13.1, §13.2).
UnitTestDesign.pairs_excursion — Function
Deprecated alias of excursions at distance 2: call excursions(input...; distance = 2).
pairs_excursion(input...; kwargs...)It takes the keywords of excursions, and the 0.4 n_way as the distance. It warns through Base.depwarn and will be removed in the next breaking release (contract §13.1, §13.2).
UnitTestDesign.triples_excursion — Function
Deprecated alias of excursions at distance 3: call excursions(input...; distance = 3).
triples_excursion(input...; kwargs...)It takes the keywords of excursions, and the 0.4 n_way as the distance. It warns through Base.depwarn and will be removed in the next breaking release (contract §13.1, §13.2).