Skip to content

Process setup

How do you define hiring evaluation criteria?

Forty applications arrive on one posting, three people read the list, and each of them pushes a different candidate forward. Nobody is being careless: as long as the criteria are unwritten, everyone reads the list they carry in their own head. This guide covers writing the criteria down, weighting them, deciding which one is a precondition and what a score means in each range, in five steps.

100

The sum of the weights. A posting is not saved until it is exactly 100, and the save button never opens

55 / 35 / 20

Default values of the strong, review and weak thresholds, stored per posting

4 items

The most rubric items one criterion can produce; the rubric is generated once per posting

12 minute readUpdated: August 2, 2026

Quick answer

Evaluation criteria are the measures that tell candidates apart on one posting and can be verified from a document. Setting them up is three jobs: writing the criteria item by item with a percentage weight on each, bringing the total weight to 100, and deciding which score range counts as strong and which as weak. Measures that call for interpretation go to the interview, not to the criterion list.

What an evaluation criterion is, and how it differs from the posting text

An evaluation criterion is a single measure you look at while reading an application. It carries two conditions at once: it tells candidates apart, and it can be verified from a document you already hold. "Join a dynamic team" in the posting text is the pitch; "at least 3 years of PostgreSQL experience" is a criterion, because it gives two candidates different answers and has something to look for in a CV.

Writing the criteria down is not enough, you also have to say which one matters more. A list without weights says every item is equal; in most roles one item outweighs the other three put together. In the panel every requirement row has its own percentage and its own "Required" switch, and the two work independently of each other.

The criteria you write turn up in three places: as the weight distribution on the posting form, as rubric items on the saved posting, and as decision bands once a score exists. This guide sets the three up in that order.

  • Criterion: written item by item in the "Requirements & Weights" section of the posting form; every row gets a percentage and a "Required" switch.
  • Rubric: the criteria are broken into check items once per posting, stored, and applied identically to every candidate.
  • Decision bands: which range of the 0-100 score counts as strong, review or weak is stored per posting.
Cards laid out side by side on a paper background, one of them with a blue strip along its edge.
A criterion list is not a wish list, it is a budget with fixed shares.

Step 1

Write criteria that can be verified from a document

Add them line by line in the Requirements section of the posting form; a line that cannot be verified is flagged as you type.

Open the Jobs area in the panel, go to the Postings tab and open the posting. The note on the "Requirements & Weights" section sums the work up: write criteria that tell candidates apart and can be evaluated from a CV, and make the weights add up to 100.

The moment you type a line, a deterministic classifier reads it. If the criterion cannot be verified from a CV, the row turns amber and the reason appears underneath it under the heading "AI cannot analyze this". Those reasons fall into four headings and their wording comes from a single dictionary, so no screen in the panel words them differently.

The warning is not decoration, it is the very triage used in scoring: the indicator you see while typing and the way a candidate is scored hang on the same classifier. A flagged row gets a "Suggest a fix with AI" button; one request returns a rewrite that can be evaluated from a CV, and "Apply this suggestion" puts it in place of your text.

  • Working pattern and availability ("three days a week", "in the office"): cannot be verified from a CV, this is a condition.
  • Cognitive abilities (analytical thinking, problem solving): cannot be verified objectively.
  • Soft skills (communication, teamwork, motivation): cannot be verified objectively.
  • General: the criterion cannot be verified from the CV or the application form, it belongs in an interview.

Step 2

Spread the weights and bring the total to 100

Give every row a percentage; the total badge counts live and Save stays closed until the total is 100.

Every requirement row has a narrow percentage field on its right. As you type, the badge under the list counts the total live and shows it as "Total: X / 100". If the total is not 100, or one row is left without a weight, both the badge and the error text beside it turn destructive.

The rule is enforced on both sides. The Save button never opens on the client and its reason stays printed under the form: "Requirement weights must sum to 100 before saving." The server runs the same schema again at write time, so there is no way to slip past the rule from the client.

One question is enough while you divide the budget: when I cannot choose between two candidates, which item do I look at. That item takes the highest percentage. Rather than splitting the rest evenly, leave clear gaps; 40 / 30 / 20 / 10 produces a ranking, 25 / 25 / 25 / 25 does not.

Two people sitting side by side at a table, looking at a single sheet and pointing at one line.
Two people settle the weights in ten minutes; left to a list, the argument runs for weeks.

Step 3

Decide which criterion is a precondition

The "Required" switch is independent of the weight and is not an elimination button, it is a cap placed on the score.

Every criterion row carries its own "Required" switch, and that flag works independently of the weight. A criterion weighted 15 can be required while one weighted 40 is not. Weight answers "how important is this", required answers "can we do without it".

The required flag eliminates nobody. An unmet required criterion puts a cap on the candidate score; the default cap is 54. The candidate stays on the list, the score cannot climb above the cap, and the call is still yours. GoTeam eliminates no candidate on its own: AI produces suggestions, and the status is always changed by a person.

To become a precondition, a criterion has to be measurable. Two routes are accepted: the text parses into the "at least N years of <technology>" pattern, or the category of the criterion is years of experience, degree, language level or grade average. A free form skill statement ("knows React") cannot be a precondition, because an elimination gate is not built on interpretation.

  • The years calculation is not asked of the model, it happens in code: 12 of 12 hard edge cases were decided correctly and no wrongful elimination came out.
  • A gate returning "unknown" does not mean "closed"; an unclear item leaves both the numerator and the denominator, and the candidate is not penalised.
  • Field of study or degree criteria ("graduate of ... Engineering") are never promoted to required on their own; a candidate from another field does not hit the cap over a diploma line alone.
  • A university year criterion ("must be a third or fourth year student") counts as an eligibility condition: it is hard even when it is off by default, and it is verified from the application form declaration rather than the CV.

Step 4

Set the decision bands for this posting

The strong, review and weak thresholds are stored per posting and have to be strictly decreasing.

The "Scoring Bands" section of the same form has three numeric fields: "Strong threshold", "Review threshold" and "Weak threshold". All three are percentages and their defaults are 55, 35, 20. Which band a score is read in follows from those three numbers; a score under the weak threshold counts as "missing".

The order cannot break: the strong threshold has to be strictly greater than the review threshold, and the review threshold strictly greater than the weak one. Not even equality is allowed, "strong 50, review 50" will not save. When the rule breaks, a red error appears and Save stays closed.

Band labels come from a single source in the product: strong, review, weak and missing. The engine, the API and the interface use the same dictionary, and no screen derives a verdict of its own from a raw number. A score that was never calculated also sits in the missing band, but that is not the same as a zero: the absence of evidence is not held against the candidate.

A long strip stretched across a paper background, split into three parts, the middle one cobalt blue.
The band is where a score turns into a decision; you set its thresholds per posting.

Step 5

Generate the rubric and approve the items one by one

Criteria are broken into check items; every item is approved together with two example pieces of evidence.

The rubric section is only drawn on a saved posting, because generation runs off the stored criterion list of that posting and the resulting version is bound to it. When you ask for "Generate rubric", the criteria are broken into check items: at most four items per criterion, all in a single call and without a single candidate being seen. The output is stored and applied identically to everyone who applies.

Approval happens item by item. For each item you are shown one example that meets it and one that does not, and you approve after seeing both. A bulk approve button is deliberately absent; the screen says so in its own note. The badges on an item read "Elimination gate", "Usage condition" and "Weight N", the last one stating its relative importance inside the rubric.

Criteria that cannot be verified from a CV never enter the rubric; those findings are listed in a separate warning box ("N criteria that cannot be verified from a CV were left out of the rubric"). If no gate item could be produced for a required criterion, the system does not turn some random item into the gate, it raises the missing gate as a validation finding on the approval screen.

Two colleagues at a high table reading the items on a page one by one, following them with a finger.
No bulk approval: every item is read together with two example pieces of evidence.

What separates a good criterion from a bad one

The most common mistake while writing criteria is copying a wish from the posting text into the criterion list. "Team player", "eager to learn" and "three days a week in the office" belong to the pitch. Since none of them can be verified from a document, they have no counterpart in evaluation and they burn weight budget for nothing.

The second mistake is subtler: the criterion is verifiable but written in a way that cannot be measured. The system does not reject such a row and it still takes its weight; it only trips the measurability test the moment you try to build a precondition on it.

Frequently written criteria and how the panel responds
Criterion as writtenHow the system respondsWrite instead
Being a team playerSoft skill: the row is flagged, no rubric item is producedTake it off the criterion list, move it to an interview question
Analytical thinking abilityCognitive skill: flagged the same wayAssess it in the interview
Able to be in the office 3 days a weekWorking pattern: this is a condition, not a requirementAdd it to the application form as a question of your own
5 years of backend development experienceSaved, but the sentence names no recognised technologyAt least 5 years of Node.js experience
2 years React, 1 year NodeTwo durations in one sentence: which duration belongs to which technology is unclearWrite two separate criterion rows
At most 2 years of experienceAn upper bound flips the condition and parsing stopsState the upper bound in the posting text, not on the criterion row

Verifiability: which document is the criterion looking at

A criterion being verifiable means its answer is written down in a document you hold. In candidate intake that document is the CV; the questions you add to the application form are the second source. Everything outside those two belongs to the interview.

The distinction has consequences: a criterion that cannot be verified produces no rubric item, so the percentage you set aside for it matches no item at all. Shortening the list is a gain here, not a loss. Four verifiable criteria produce a sharper ranking than a mixed list of eight.

  • Verified from the CV: years of experience, tools used, degree level, language level, grade average.
  • Verified from the application form: declared facts such as university year, work permit or start date.
  • Belongs to the interview: communication, team fit, motivation, approach to problem solving.

Measurability: can the criterion be a precondition

Measurability only comes into play when you build a precondition. The moment you mark a criterion "Required", rubric validation looks for a complete gate item for it and can raise four separate findings: no gate at all, more than one gate, an unmeasurable gate, an empty item question.

The deterministic parser is conservative on purpose. If the text holds more than one duration, if the technology name is not recognised in the canonical skill dictionary, if several different technologies appear, or if there is an upper bound phrase, it stops parsing. That is not a shortcoming, it is a brake fitted so that a wrong gate is never set up quietly.

Company wide settings and per posting settings

Part of criterion setup concerns one posting only, part of it changes every posting in the company at once. Mixing the two is the most common setup mistake: changing the company wide preset because a single posting should be "a bit stricter" shifts the behaviour of every open posting.

There are three company wide presets: Balanced (the default), Strict and Lenient. Strict pulls the cap for an unmet required criterion down to 40 and keeps the default requirement behaviour on. Lenient lifts the cap to 60 and turns the default requirement off. This screen is separate from the per posting decision bands.

A practical rule: if the ranking on one posting is not coming out the way you expected, look first at the criteria, the weights and the bands of that posting. Touch the company wide setting only when you see the same drift on several postings.

Where each setting lives
SettingWhereScope
Criterion text, weight, the "Required" switchPosting form, Requirements & WeightsThat posting only
Strong / Review / Weak thresholdsPosting form, Scoring BandsThat posting only
Rubric items and versionPosition Rubric section on a saved postingThat posting only
Required criterion cap (default 54)Settings, AI scoring screenEvery posting in the company
Default requirement behaviourSettings, Balanced / Strict / Lenient presetEvery posting in the company

What changing a criterion later costs you

If calibrating criteria were expensive nobody would do it, so the price of this path was kept low on purpose. One candidate and posting pair is charged only once; no credit is spent when a rescoring follows a change to the criterion text, a weight or a band. Stored coverage calls are re-added locally and no new model request is needed.

The change is not silent either. On a posting that already has analysed applications, trying to save an edit to a criterion text, its weight or its required flag opens a confirmation window that tells you with a number how many analyses will be affected. If nothing changed, or no analysis is affected, the window never opens.

Being able to look back is part of the setup too: every score record keeps the version, the model and the gate rationale, so why a candidate ended up in that band can be shown afterwards. The system does not learn from past hiring outcomes; who was invited or hired never enters scoring, which means the ranking does not drift on its own as long as you leave the criteria alone.

  • Writing criteria, weighting, decision bands and rubric approval are locked on no plan; the screens ask for posting view and update permissions.
  • The path that drafts criteria and bands from a free text (AI posting generation) is open on Professional, Enterprise and the 14 day trial; it is closed on the Mini and Starter plans. While it is closed you write the criteria by hand and the rest works the same.
  • Rubric generation spends one AI evaluation credit per posting; the requirement check and the re-adding that follows a band or weight change are free.
  • While the AI analysis switch in company settings is off, the scoring settings and the re-analysis endpoints close, and a generated rubric has no counterpart in scoring.
Stacked paper folders with the blue one in the middle standing out among the others.
Changing a criterion erases no history; which version produced a score stays on the record.

Frequently asked questions

Because a weight is not a priority, it is a share of a budget: a criterion only says something about the score while the total is fixed. If the total were free to land anywhere between 90 and 110, the same percentage would mean two different things on two postings. Both sides enforce the rule: the save button stays closed in the browser, and the server validates it again with the schema at write time.

No. Required is the switch on the row, and when it is not met it caps the score of that candidate; the default cap is 54. An elimination gate is an item inside the rubric and it can only be built on a measurable criterion. Mark an unmeasurable criterion as required and rubric validation returns the finding that a gate item has to be measurable. Neither of them removes a candidate from the list, the decision is still yours.

You can, the system does not delete the row. But a criterion like that cannot be verified objectively from a CV, so it produces no rubric item and the percentage you spent on it has no counterpart in the scoring. The row is flagged with an amber warning as you write it. Its proper place is not the criteria list but the interview, where the same competency is assessed through questions and answers.

Yes, all three thresholds are set per posting in the Scoring Bands section of the posting form. The defaults are strong 55, review 35 and weak 20. The only constraint is the order: strong has to be strictly above the review threshold and review strictly above the weak threshold, equality is rejected. If you pull the strong threshold below 55, keep in mind that the cap for a missed required criterion is 54.

The rubric is generated once per posting, in a single call and without seeing a single candidate; the output is stored and applied identically to everyone who applies. If you choose to regenerate it, every click spends one AI evaluation credit. Editing the item texts by hand and saving them as a new version never reaches the model and costs nothing.

Evaluations are not deleted, they are re-added, and that is free: one candidate and one posting pair is charged only once. Before you save, a confirmation window counts how many analyses the change will touch, so nothing happens silently. The ranking can move; because every score record keeps the version, the model and the reason behind the gate, you can always show after the fact why a candidate sits in a given band.

Let us set the criteria up on your own posting

The weight table, the decision bands and rubric approval are waiting in the panel. Request a demo and we will walk through the setup together.

Contact Us