Critical thinking test guide

Critical thinking tests and the five sections

A critical thinking test is not one task. It is five, each with its own verdict rule, and a rule carried from one section into the next produces a confident answer to the wrong question.

5 drill areasreview queue

146practice questions
6free worked samples
1provider families

The guide, the samples and the diagnostic are free. Full access is one package covering every provider and test type.

Reviewed 19 August 2026 by the Cognivy editorial team

The critical thinking assessment inside Cognivy's launch scope is Pearson's Watson-Glaser Critical Thinking Appraisal, and it behaves unlike other reasoning tests. The same passage can support one verdict in the inference section and a different one in the interpretation section, because the two sections apply different thresholds. This guide sets out what Pearson publishes about each section, the threshold each one uses, and six free practice questions solved against the rules.

Questions
40, as published by Pearson
Time limit
30 to 40 minutes, timed
Answer format
Fixed verdict sets that change by section, from the five-point inference scale to assumption made or not made
Sections
Five: drawing inferences, recognising assumptions, deducing, interpreting and evaluating arguments
Scoring
Number correct with percentile, T-score, stanine and sten. No pass mark is published

The skill

What is a critical thinking test?

A critical thinking test measures how well you handle evidence: what it supports, what it takes for granted, and how strongly. The dominant product in graduate and professional hiring is Pearson's Watson-Glaser Critical Thinking Appraisal, which is built from five separate exercises rather than one repeated question type.

Pearson describes the five as drawing inferences, recognising assumptions, deducing, interpreting, and evaluating arguments. Each has a short passage or statement and a small fixed response set, and each response set means something different. That is the structural fact worth learning before any content.

The material is deliberately ordinary — policy, business and public-interest passages — because the test is not measuring subject knowledge. It is measuring whether you can hold a position you personally disagree with, and refuse a conclusion you personally believe, on the evidence in front of you.

Not a general verbal reasoning test

A verbal reasoning test usually asks one question repeatedly: is this statement true, false or impossible to say from the passage. A critical thinking appraisal switches the question five times, and the verdict words change meaning as it does. If your invitation names Watson-Glaser, prepare the five sections separately. If it names a verbal or reading assessment, the verbal reasoning guide is the closer match.

Try it before anything else

Free critical thinking practice questions

Six original questions: one for each of the five sections, in the order Pearson lists them, then a last one on naming the flaw in an argument. Commit to a verdict before opening the solution — each solution applies the section's rule and names why every other option fails.

Sample question 1 — an inference itemQuestion 1 of 6

Inference items give a passage and a proposed inference, and ask how probable the inference is on the passage alone. The response set runs true, probably true, insufficient data, probably false, false.

Passage: A city council reports that 1,240 of the 2,000 residents who responded to its consultation supported extending the cycle lane. The consultation was open online for four weeks and was publicised in the council newsletter.

Inference: a majority of the city's residents support extending the cycle lane. How should it be rated?

  • A. True
  • B. Probably true
  • C. Insufficient data
  • D. Probably false
  • E. False
Show the worked solution

Answer: C — Insufficient data.

  1. Identify what the passage actually counts: 1,240 of 2,000 respondents, which is a majority of the people who replied.
  2. Identify what the inference claims: a majority of the city's residents, which is a different and much larger group.
  3. The passage gives no population figure and no evidence that respondents resemble residents, and a self-selected online consultation is not stated to be representative.
  4. Because the passage neither supports nor contradicts the claim about residents, the correct rating is insufficient data.

Why the other options are there. True treats 1,240 out of 2,000 as if respondents were the population, a step the passage never licenses. Probably true assumes representativeness that the passage never offers. Probably false and false both require evidence pointing the other way — that most residents oppose it — and the passage contains none, so neither is available either.

Sample question 2 — a recognition of assumptions itemQuestion 2 of 6

Assumption items give a statement and a proposed assumption, and ask whether the statement takes it for granted. An assumption is unstated by definition, so the test is whether the statement still stands without it.

Statement: We will cut delivery times by moving all parcels to the new sorting hub.

Proposed assumption: the new hub can process the parcel volume currently handled elsewhere. Is this assumption made?

  • A. Assumption made
  • B. Assumption not made
Show the worked solution

Answer: A — Assumption made.

  1. Restate the plan as a mechanism: moving all parcels to the hub is expected to reduce delivery times.
  2. Apply the negation check. Suppose the hub cannot process the volume currently handled elsewhere.
  3. Under that supposition, moving all parcels there would create a backlog and delivery times would not fall, so the statement's own plan would fail.
  4. Because the statement collapses without it, the capacity of the hub is taken for granted, and the assumption is made.

Why the other options are there. Assumption not made is chosen by candidates who look for the words in the statement and cannot find them. That is the wrong test: an assumption is never written down, or it would be a premise. The reliable check is negation, because it tests whether the statement depends on the assumption rather than whether the words appear.

Sample question 3 — deductionQuestion 3 of 6

Deduction questions give one or two premises and a proposed conclusion, and ask whether the conclusion follows of necessity. Pearson describes the section as determining whether conclusions follow logically from the given information, so plausible is not enough and outside knowledge does not count.

Premises: All of the firm's senior engineers have completed the security induction. Some employees who have completed the security induction work remotely.

Conclusion: some of the firm's senior engineers work remotely. Does the conclusion follow?

  • A. Conclusion follows
  • B. Conclusion does not follow
Show the worked solution

Answer: B — Conclusion does not follow.

  1. Map the groups. Every senior engineer sits inside the group that completed the induction, and some members of that group work remotely.
  2. Ask what the premises force. The remote workers are somewhere among the induction-completers, but nothing places any of them among the senior engineers.
  3. Build the counterexample: every remote induction-completer could be a warehouse supervisor, with all the senior engineers on site. Both premises hold and the conclusion is false.
  4. A conclusion that a consistent counterexample defeats is not forced, so it does not follow, however likely it is to be true in fact.

Why the other options are there. Conclusion follows is reached on plausibility: both premises pass through the same induction group, which invites the mind to merge them, and engineers working remotely sounds entirely normal. But the some in the second premise carries no information about which induction-completers it picks out, and real-world likelihood is exactly what this section tells you to set aside. The test is whether the premises force the conclusion, not whether it is likely to be true.

Sample question 4 — interpretationQuestion 4 of 6

Interpretation questions give a short passage and a proposed conclusion, and ask whether the conclusion follows beyond reasonable doubt. Pearson describes the section as weighing the evidence and deciding whether conclusions based on data are warranted, which is a lower bar than deduction's: warranted, not forced.

Passage: A firm's annual staff survey is anonymous and was completed by all 640 employees in each of the last two years. This year, 74% said they would recommend the firm as a place to work. Last year, 61% said so.

Conclusion: willingness among the firm's employees to recommend it as a place to work was higher this year than last. Does the conclusion follow?

  • A. Conclusion follows
  • B. Conclusion does not follow
Show the worked solution

Answer: A — Conclusion follows.

  1. Fix what was measured: the share of the workforce saying they would recommend the firm rose from 61% to 74%, and every employee answered in both years, so no sampling gap opens between the surveyed group and the workforce.
  2. Name the step the conclusion takes beyond the arithmetic: it reads what people said as tracking what they were willing to do, which no survey can strictly guarantee.
  3. Apply this section's threshold. Interpretation does not ask whether the conclusion is forced, only whether the passage warrants it beyond reasonable doubt.
  4. The one doubt left is that anonymous answers misstated willingness, and did so differently across the two years. That is a theoretical doubt rather than a reasonable one, so the conclusion follows.

Why the other options are there. Conclusion does not follow is deduction's verdict imported into the wrong section: sincerity is not logically guaranteed, so the conclusion is not forced, but this section never asked for forced. Marking every unforced conclusion as not following fails interpretation as surely as accepting every plausible one fails deduction. Reading which exercise you are in is part of the task.

Sample question 5 — evaluation of argumentsQuestion 5 of 6

Argument questions give a question of policy and an argument on one side, and ask whether the argument is strong or weak. Pearson describes the section as evaluating the strength and relevance of arguments with respect to a particular question, and being true is not enough: a strong argument must be both directly related to the question and important.

Question at issue: should the city ban private cars from the city centre on weekdays? Argument: no — drivers who are used to crossing the centre would have to learn new routes.

Is the argument strong or weak?

  • A. Argument strong
  • B. Argument weak
Show the worked solution

Answer: B — Argument weak.

  1. Test relevance first. The argument describes a direct effect of the ban on drivers, so it addresses this question rather than a different one.
  2. Test importance next. Set against a decision about how a city centre works every weekday, having to learn new routes is a minor, one-off inconvenience.
  3. An argument must clear both bars, and this one clears only relevance, so it is weak.
  4. Notice that the verdict never involved whether a weekday ban is a good idea. The section scores the argument, not your position on the question.

Why the other options are there. Argument strong is reached two ways, and both are errors this section is built to catch. Some candidates rate it strong because it is undeniably true that drivers would need new routes, but being true is not enough; only relevance and importance make strength. Others rate it strong because they oppose the ban and welcome any argument on that side, which is agreement rather than evaluation.

Sample question 6 — naming the flawQuestion 6 of 6

The last format asks you to name what has gone wrong in an argument rather than to grade a verdict. Recognising the standard failure patterns is what makes the deduction and arguments sections fast, because a flaw you can name is a flaw you stop re-deriving from scratch.

Argument: if the product launch is delayed, the marketing budget will be underspent this quarter. The launch is going ahead on schedule. Therefore the budget will not be underspent.

Which description names the flaw in this argument?

  • A. It treats a delay as the only possible route to an underspent budget, though the premise offers it only as one
  • B. It assumes from the outset the very conclusion it sets out to prove
  • C. It generalises about every quarter's budget from the evidence of a single launch
  • D. It infers a connection from two events that have merely coincided in the past
Show the worked solution

Answer: A — the argument treats a sufficient condition as a necessary one.

  1. Write out the form: if delay, then underspend. No delay. Therefore no underspend.
  2. The first premise says a delay would be enough to underspend the budget. It does not say a delay is required for that.
  3. So ruling out the delay closes one route to an underspent budget, not all of them: a cancelled campaign or an unfilled vacancy could underspend it with the launch exactly on time.
  4. The conclusion therefore goes beyond the premises, and option A describes how: one sufficient route has been treated as the only route.

Why the other options are there. Option B fails because the conclusion, that the budget will not be underspent, appears in neither premise, so nothing is assumed in advance. Option C fails because the argument never generalises; the link between this launch and this quarter's budget is granted as a premise, not inferred from a run of cases. Option D fails because no past co-occurrence is cited anywhere: the argument works from a stated conditional, and its error lies in the inference drawn from it, not in the evidence offered for it.

Sit a timed free set

A short critical thinking set at the real per-question pace, scored server-side like the full product, with a worked solution after every answer.

Start the timed free set

Draws on your 20 free practice questions. No account needed.

What employers are looking at

What do critical thinking tests measure?

Pearson describes the Watson-Glaser appraisal as a quick, consistent and accurate measure of the ability to analyse, reason, interpret and draw logical conclusions, and presents its cognitive assessments as a way to select candidates with high potential to perform and develop, across roles it lists as including manufacturing, logistics, legal, sales and communication.

Two demands run through all five sections. Threshold discipline: can you tell the difference between an inference that is probably true and one the passage cannot settle. Detachment: can you evaluate an argument for a position you dislike on its relevance and importance rather than on whether you agree with its conclusion.

Reading your result

How critical thinking tests are scored

A raw count out of 40 is only the first line of what Pearson reports. Everything after it is comparative, which changes what a good result means and what preparation should optimise for.

  • The report carries five score types, not one. Pearson reports Watson-Glaser results as a number correct alongside percentile, T-score, stanine and sten scores, in Profile, Development or Interview report formats. Every figure after the number correct places you against other people rather than against a fixed standard.
  • A percentile is a position in a norm group. A percentile compares you with a comparison group rather than with a threshold, which is why a raw total on its own does not tell you much. The same performance can land at different percentiles against different norm groups, so a number without its group means little.
  • No pass mark is published. Pearson publishes score types rather than a threshold. Any cut-off is set by the employer for the role, so the only benchmark worth working to is one the recruiter has actually disclosed, and asking is the only way to learn it.
  • Two candidates may not sit identical questions. Pearson describes delivery as item-banked with instant results, which means your set of questions can differ from another candidate's. Preparing the five verdict rules transfers to whatever is drawn; memorising specific practice questions does not.
  • One overall figure hides where the marks went. The five sections make different demands, so when you practise, record accuracy per section rather than as a single number. The repair for a weak inference section, re-learning the thresholds of the five-point scale, has nothing in common with the repair for a weak arguments section.

Question formats

The five sections and their verdict rules

Pearson publishes what each of the five sections asks. The response set changes with the section, and so does the standard of proof.

  • Drawing inferences. Pearson describes this as rating the probability of the truth of inferences based on the information given. The response set is graded — true, probably true, insufficient data, probably false, false — so the section is about degree of support, not a yes or no.
  • Recognition of assumptions. Pearson describes this as identifying unstated assumptions or presuppositions underlying given statements. An assumption is by definition not written down; the test is whether the statement collapses without it.
  • Deduction. Pearson describes this as determining whether conclusions follow logically from the given information or data. This is the one section with a strict standard: a conclusion follows only if it is forced, and personal knowledge is irrelevant.
  • Interpretation. Pearson describes this as weighing the evidence and deciding if generalisations or conclusions based on data are warranted. This is the softer of the two standards: deduction accepts a conclusion only when the information forces it, while interpretation asks whether the passage warrants it beyond reasonable doubt. A conclusion that is not forced can still be warranted, which is why the same statement can fail deduction and pass interpretation.
  • Evaluation of arguments. Pearson describes this as evaluating the strength and relevance of arguments with respect to a particular question or issue. A strong argument must be both directly related to the question and important; being true is not enough.
  • Passages that repeat across sections. The same or similar material can appear under more than one section heading. Reading which exercise you are in is part of the task, not a formality to skim.
Why the same passage gives two answers

A passage reports that 1,240 of 2,000 consultation respondents backed a proposal. In the inference section, the statement that most residents back the proposal is not something the passage settles, because respondents are not residents. In the assumptions section, a plan that says the proposal is popular because the consultation supported it does take for granted that respondents represent residents. Same passage, same gap, two different jobs: rating support, and naming what is being presumed.

Launch coverage

Which providers use critical thinking tests?

One launch family owns this format outright. The others assess neighbouring skills under different names.

  • Watson-Glaser Critical Thinking Appraisal. Pearson lists the appraisal at 40 items and 30 to 40 minutes timed, delivered with item-banked questions and instant results, and suitable both for unsupervised screening and for supervised online completion. Pearson names the five measured components as drawing inferences, recognising assumptions, deducing, interpreting and evaluating arguments, and offers Profile, Development and Interview reports.

Provider names identify the assessment format. Cognivy is independent and is not affiliated with or endorsed by these providers.

What varies, and what does not

Timing, scoring and what the report shows

Only one provider in the launch scope publishes a critical thinking appraisal, so the published figures are specific rather than a range.

  • 40 items, 30 to 40 minutes, timed. Pearson lists those figures for the appraisal as a whole. It does not publish a separate limit for any one of the five sections, so the section split is yours to manage.
  • Item-banked, with instant results. Pearson describes the delivery as supporting remote completion, item-banked questions and instant results, which means two candidates may not see identical items.
  • Supervised and unsupervised are both supported. Pearson states the appraisal is suitable for unsupervised screening and for online completion in a supervised environment. Which one applies to you changes what you may have to hand, so read your own instructions.
  • No pass mark is published. Pearson publishes score types rather than a threshold. Any cut-off is set by the employer for the role, and it is worth asking the recruiter rather than assuming a figure.

A repeatable approach

A method that survives the clock

The reliable approach is procedural. Name the section, then apply that section's rule and no other.

  1. 1

    Name the section before you read the passage. Inference, assumption, deduction, interpretation or arguments. The single largest source of error in this test is answering a good question that the current section did not ask.

  2. 2

    Say the threshold out loud in your head. Inference asks how probable. Deduction asks whether it is forced. Interpretation asks whether the evidence warrants it beyond reasonable doubt, which is a lower bar than being forced. Arguments asks whether it is both relevant and important. Assumptions asks whether the statement needs it.

  3. 3

    Underline the scope words. All, most, some, only, always. A scope error is what turns a supported claim into an unsupported one: the passage supports a claim about respondents and the option makes a claim about everyone.

  4. 4

    Answer from the passage, then check for your own view. Ask explicitly whether you would have given the same verdict for the opposite conclusion. If not, you have answered on agreement rather than evidence.

  5. 5

    Commit and move. With 40 items and 30 to 40 minutes, an item that has become a debate with yourself is an item to answer and leave. Re-reading a passage a third time rarely changes the verdict and always costs the next question.

What slows progress

Common mistakes, and the fix for each

  • Carrying one section's rule into another. Fix: name the section before each item. A conclusion that is not deductively forced may still be warranted in the interpretation section, and an inference the passage cannot settle is not the same as one it contradicts.
  • Answering on agreement. Fix: ask whether you would give the same verdict if the conclusion pointed the other way. The appraisal deliberately uses topics people hold views about, because that is where the measurement lives.
  • Looking for an assumption in the text. Fix: use the negation test. Assume the proposed assumption is false, and see whether the statement still works. If it does not, the assumption is being made.
  • Treating relevance as strength. Fix: a strong argument must be relevant and important. Pearson defines the section as evaluating strength and relevance with respect to the particular question, so an argument that is true but addresses only a narrow case is weak.
  • Ignoring scope words. Fix: underline all, most, some and only before deciding. The difference between what the passage supports and what the option claims is usually one of these words.

You have seen the method. Employers set the full battery.Every provider, every track, a worked solution on every question.

Get full access

In practice

How does the five-point inference scale work?

The inference section's response set runs true, probably true, insufficient data, probably false, false, and Pearson describes the task as rating the probability of the truth of inferences from the given information. It is a scale of support, not a yes or no: the outer verdicts are for inferences the passage settles outright, the probably verdicts for inferences it makes likely without settling, and insufficient data for claims it cannot decide in either direction.

The middle verdict does the most work and attracts the most errors. Insufficient data is not a refuge for hard questions; it is a positive finding that the passage neither supports nor contradicts the claim. Sample question 1 turns on it: a majority of respondents is what the passage counts, a majority of residents is what the inference claims, and nothing connects the two groups, so the claim cannot be rated in either direction.

The other discipline the scale imposes is that the false end needs evidence, not absence. Probably false and false both require the passage to point against the inference. A claim the passage merely fails to support is insufficient data, never probably false, and demoting an unsupported claim to the false end is as much an error as promoting it to true.

In practice

What counts as an assumption an argument actually makes?

An assumption is something a statement takes for granted without saying, and Pearson describes the section as identifying unstated assumptions or presuppositions underlying given statements. That definition rules out the most natural approach, which is scanning the statement for the assumption's words: if the words were there, it would be a premise rather than an assumption, so the search proves nothing either way.

The reliable test is negation. Suppose the proposed assumption is false, then re-run the statement and see whether it still works. The sorting-hub plan in sample question 2 collapses the moment the hub cannot process the volume, so the hub's capacity is being taken for granted even though no sentence mentions it. Collapse under negation is what assumption made means.

The boundary runs the other way as well. A proposed assumption can be sensible, relevant and even true without being one the statement needs, and if the statement survives its negation intact, the verdict is assumption not made. The section asks what this statement leans on, not what a careful person would also like to be true.

In practice

Why can the same conclusion pass interpretation and fail deduction?

Because the two sections apply different standards to the same kind of question. Pearson describes deduction as determining whether conclusions follow logically from the given information, and interpretation as weighing the evidence and deciding whether conclusions based on data are warranted. In practice that is the gap between forced and warranted beyond reasonable doubt: deduction accepts a conclusion only when no counterexample fits the premises, while interpretation accepts one that the passage leaves no reasonable ground to doubt.

Sample questions 3 and 4 sit either side of that line deliberately. The remote-working conclusion is entirely plausible and still fails deduction, because a consistent counterexample exists. The survey conclusion is not strictly forced, since answers could in theory misstate willingness, and still passes interpretation, because the doubt that remains is theoretical rather than reasonable.

This is why naming the section is the method's first step. Carrying deduction's standard into interpretation marks warranted conclusions as not following, and carrying interpretation's tolerance into deduction waves through merely plausible ones. Both produce the same characteristic error: a careful answer to a question the current section did not ask.

A clear route

How to prepare in the days you have

The five sections need separate work, because mixing them before each verdict rule is secure leaves you unable to tell which rule you actually applied. This order fits a short run-up.

  1. 1

    First, learn the five verdict rules cold. Write out what each response option means in each section until you can recite them. Everything else in this test depends on it, and unlike the rest of the appraisal it is a fixed list rather than a skill, so it is the one part of preparation that has an end point.

  2. 2

    Next, practise one section at a time, untimed. Inference and assumptions first, because they are the least like other reasoning tests you may have sat. Deduction transfers most easily from the deductive reasoning guide.

  3. 3

    Review by rejected option, not by score. For every item, state why each wrong verdict fails. Understanding why probably true was wrong teaches the threshold; knowing the key does not.

  4. 4

    Then mix the sections. Only once each rule is stable. Mixed practice trains the section-switch, which is the specific skill the full appraisal adds.

  5. 5

    Finish with a full timed rehearsal. Pearson lists 40 items in 30 to 40 minutes, so rehearse at roughly that rate. Pearson also publishes a free Critical Thinking (Watson Glaser) practice test, which it describes as intended to show how the test works and what the questions in each subtest are like.

Cognivy uses your assessment date to choose the route rather than asking you to predict a study schedule. When you sit down to practise, you choose the session length that fits that day.

Pace

Getting faster without losing accuracy

At around a minute an item across five different tasks, pace comes from deciding cleanly rather than reading quickly.

  • Read the question stem first. Knowing which section you are in tells you what to look for in the passage, which usually means one pass rather than two.
  • Use the negation test as a reflex. On assumptions it produces a definite answer rather than a feeling: either the statement collapses without the assumption or it does not.
  • Decide on the scope word. When an option generalises past what the passage counted, the verdict is usually settled by that alone, with no further analysis needed.
  • Set a per-item ceiling. Around one minute is the working average implied by Pearson's published length. Passing that ceiling is a signal to answer and move rather than to read again.

Direct answers

Critical thinking test FAQs

What is a critical thinking test?

It is an assessment of how you handle evidence: what a passage supports, what it takes for granted, and how strongly. In employer hiring the usual product is Pearson's Watson-Glaser Critical Thinking Appraisal, which is built from five separate exercises rather than one repeated question type.

What are the five sections of the Watson-Glaser test?

Pearson names them as drawing inferences, recognising assumptions, deducing, interpreting, and evaluating arguments. Each has its own response set and its own standard of proof, which is why the same passage can produce different verdicts in different sections.

How long is the Watson-Glaser test?

Pearson lists the Watson-Glaser Critical Thinking Appraisal at 40 items and 30 to 40 minutes, timed. It does not publish a separate limit for each of the five sections, so managing the split between them is part of the task. Use the time stated in your own invitation.

What is a good Watson-Glaser score?

Pearson publishes score types rather than a pass mark: number correct, percentile, T-score, stanine and sten. A percentile compares you with a norm group, and any cut-off is set by the employer for the role. Ask the recruiter whether a target has been disclosed.

What is the difference between critical thinking and verbal reasoning tests?

A verbal reasoning test normally asks one question repeatedly against a passage, most often true, false or cannot say. A critical thinking appraisal switches between five exercises, and the meaning of each verdict changes with the section. The reading skill overlaps; the decision rules do not.

How do you answer recognition of assumptions questions?

Use the negation test. Assume the proposed assumption is false and see whether the original statement still holds. If the statement collapses, the assumption is being made. Searching the statement for the words will not work, because an assumption is unstated by definition.

Where is the Watson-Glaser test used?

Pearson does not publish a customer list. It presents the appraisal as a general measure of the ability to analyse, reason, interpret and draw logical conclusions, within a cognitive range aimed at selecting candidates across roles it lists as including manufacturing, logistics, legal, sales and communication. If your invitation names Watson-Glaser, that is the assessment to prepare, whatever the sector.

Are Cognivy's critical thinking questions official Watson-Glaser questions?

No. Cognivy is independent and is not affiliated with or endorsed by Pearson or any other assessment provider. Every question is original material written to teach the five verdict rules and the pacing of the format.