← Guides

Situational Judgement Test Guide: Formats and Answers

A situational judgement test guide covering the four answer formats, how SJTs are scored against expert benchmarks, two worked scenarios, and the traps that lower scores.

Situational Judgement·Aptitude Test Prep Team · Sep 2, 2026 · 8 min read

A situational judgement test shows you a short workplace scenario and several possible actions, then asks you to judge them. There is no data to calculate and no pattern to spot. Your answers are compared against a benchmark set by job experts when the test was designed, and your match is reported to the employer as an overall score plus competency breakdowns such as communication, planning, analysis and working with people. Those are usually expressed as percentiles against a norm group rather than a pass mark. The format matters more than most candidates realise, because “what would you do” and “what is the most effective action” are scored differently. This guide covers the four formats, the scoring model, two worked scenarios, and the errors that cost marks.

TL;DR

  • SJTs measure judgement in social and work situations: conflict handling, prioritisation, teamwork, escalation.
  • Four answer formats exist: most and least effective, rated effectiveness, ranked order, and likely to perform.
  • Scoring compares your answers to expert “best fit” responses, then benchmarks you against a norm group.
  • Most SJTs are not strictly timed, but you should answer promptly. A common guideline is around 20 minutes for 24 questions.
  • The winning answer is usually the one that solves the problem directly and keeps the right people informed.
  • Read the role description first. The same scenario has different best answers for a graduate and a manager.

What does a situational judgement test actually measure?

SJTs assess effectiveness in situations you would face in the job. Public sector guidance from the US Office of Personnel Management describes them as measuring conflict management, interpersonal skills, problem solving, negotiation, facilitating teamwork and cultural awareness, and notes they work particularly well for managerial and leadership competencies.

Two properties explain why employers like them. Content validity is high, because the scenarios are drawn from real tasks in the role. Criterion-related validity is moderately high, meaning performance on the test relates moderately to performance on the job. Candidates also tend to see them as fair, and subgroup differences are typically moderate.

That has a practical consequence. An SJT is written around a specific competency framework. If the employer publishes its values or graduate competencies, read them before you sit the test. They are effectively the marking scheme.

Which answer formats will you see?

There are four, and they change your strategy.

Format The instruction How to approach it
Most and least effective Pick the best action and the worst Find the clear extremes first, ignore the middle
Rated responses Rate each option on a scale from counterproductive to very effective Rate each option on its own merits, use the full scale
Ranked responses Put all options in order, each rank used once Anchor the top and bottom, then place the middle
Likely to perform Say what you would most and least likely do Behavioural, so consistency across items matters

The last one is the one people get wrong. “What would you do” measures your behavioural tendency, and publishers often build consistency checks across items. “What is most effective” measures judgement of best practice. If the instruction says “would”, answering with a textbook ideal you would never actually do can read as inconsistent.

Delivery varies too: text only, video clips, or animated scenarios with computer-generated avatars. The judgement being tested is the same.

Before the worked examples, if you also face a numerical stage in the same sitting, you can try five free timed numerical questions to check your pace.

Worked scenario 1: competing deadlines

Scenario. You are a graduate analyst. A client report is due to your manager at 16:00. At 14:30 a colleague asks for urgent help with a spreadsheet error that is blocking a different client deadline at 15:00. You estimate the fix would take 40 minutes. Your manager is in a meeting until 15:30.

Rate each response from 1 (counterproductive) to 4 (very effective).

A) Drop your report and spend 40 minutes fixing the colleague’s spreadsheet. B) Tell your colleague you cannot help and continue with your report. C) Spend five minutes looking at the spreadsheet to see whether the cause is quick to identify, tell your colleague what you find, and send your manager a short message flagging the clash. D) Message your manager that you may miss 16:00 and wait for a reply before doing anything else.

Suggested ratings and reasoning.

C scores 4. It is the only response that gathers information cheaply, helps without abandoning your own commitment, and informs the person who owns the priority decision. Note that it does not ask the manager to decide before you have done anything useful.

A scores 2. Helpful in intent, but it silently trades a certain miss on your own deadline for someone else’s, without telling anyone.

B scores 2. It protects your deadline but ignores a colleague and a client with an earlier deadline, and it shares no information.

D scores 1. It stops all work and makes waiting the plan, when the manager is unavailable for an hour and both deadlines are inside that window.

The principle. Take the cheap diagnostic action, then escalate with information rather than with a question. Answers that only escalate, and answers that only act, both score below the answer that does a little of each.

Worked scenario 2: a mistake you spot in someone else’s work

Scenario. You are reviewing a slide pack that a senior colleague will present to a client tomorrow morning. You notice a figure on one slide contradicts the figure in the appendix. You are fairly confident the slide is wrong, but you were not involved in the analysis. It is 18:00.

Rank the four responses from 1 (most effective) to 4 (least effective).

A) Correct the slide yourself and send the amended pack back with the change highlighted. B) Message the senior colleague now, name the slide and the appendix page, state the discrepancy, and ask which figure is right. C) Say nothing, because it is not your analysis and you may be wrong. D) Raise it in the client meeting tomorrow if the figure comes up.

Suggested ranking. B first, A second, D third, C fourth.

Reasoning. B is fastest, precise and leaves the decision with the person who owns the number. Naming the exact slide and appendix page is what makes it effective, because a vague “something looks off” wastes the colleague’s evening. A is well intentioned but assumes you know which figure is correct, and an unrequested edit to someone else’s analysis can introduce a second error. D creates client-facing risk from an internal problem, which is worse than saying nothing quietly, but at least surfaces it. C leaves a known error in a client deliverable, which is the lowest ranked because inaction on a spotted error is the failure the item is testing.

The principle. Raise the issue early, be specific, and leave ownership where it belongs. Never let an internal error reach the client.

How are situational judgement tests scored?

Your responses are compared to the “best fit” answers agreed by subject matter experts during test design. The number of items you rate or rank in line with those experts becomes your raw score. That is then compared against a norm group of previous test takers and reported as a percentile, usually with a competency breakdown. Published guidance commonly describes a strong result as sitting around the 70th to 80th percentile, though the employer sets the sift level.

Two implications follow. First, partial credit is normal in rated and ranked formats, so a near miss still earns something. Second, there is no negative marking to fear, so answer every item.

What lowers scores most often?

  • Escalating everything. Passing every problem upward reads as low ownership.
  • Escalating nothing. Solving things silently reads as poor communication, especially where risk or a client is involved.
  • Ignoring the stated role. A team leader is expected to delegate and set priorities. A first-week graduate is not.
  • Flat ratings. Marking almost everything “effective” gives the scoring model nothing to work with. Use the full scale.
  • Overthinking a “would do” item. Trust your first read. The format is designed for prompt answers.

Many employers run an SJT alongside other stages. Our SHL general ability test guide covers the cognitive stage, the Cappfinity assessment guide covers strengths-based formats, and how to pass online assessment tests covers the logistics of a multi-stage sitting.

Frequently Asked Questions

Are situational judgement tests timed?

Often there is no hard limit, but publishers give a guideline. One common example is roughly 20 minutes for 24 questions. Prompt answers are encouraged, because the format rewards judgement rather than deliberation.

Should I answer honestly or strategically?

Answer as the strongest version of yourself in that role. Guessing at an imagined ideal you do not recognise leads to inconsistent answers across items, and “would do” formats include consistency checks.

Can you fail a situational judgement test?

You can score below the employer’s cut-off. There is rarely an absolute pass mark, because scores are reported as percentiles against a norm group.

Do SJTs use video?

Some do. Scenarios may be text, video clips, or animation with computer-generated avatars. The response formats and scoring model stay the same.

How do I prepare when there are no right answers to learn?

Learn the employer’s competency framework, practise recognising the four formats, and rehearse the escalation principle: act cheaply, inform early, be specific.

Are SJTs used at every level?

They are used from graduate entry to management. Guidance from OPM notes they are particularly effective for managerial and leadership competencies, and the scenarios scale with the level of the role.

Sources

The 40-question SHL numerical pack, 9 dollars

Launching soon.