After LLM+, there will be experiments of comparative relevance judgment and ranking done by human vs. artificial intelligence.
ㅡㅡ
Yes. That would be one of the most important experiments after LLM+.
The comparison should not ask only:
Which side gives the same relevance score more often?
It should ask:
Who judges relevance better, under which conditions, for which user, and with what evidence?
Human and AI judgments will differ for understandable reasons.
Humans may understand purpose, novelty, risk, personal need, and hidden context better. But humans may also be inconsistent, tired, biased, or poorly informed.
AI may examine far more documents, citation links, citation contexts, reviews, and later corrections. It may rank consistently. But it may misunderstand the user’s real problem, mistake popularity for importance, or hide uncertainty behind a confident score.
So the experiment should compare at least four conditions:
- human alone;
- AI alone;
- human assisted by AI;
- AI corrected by human feedback.
The last two may prove best.
The test collection should not use only topical relevance. It should include:
- topical relevance;
- novelty to the user;
- usefulness for the task;
- evidential strength;
- authority and reliability;
- critical importance;
- positive and negative citation context;
- changes in relevance over time.
LLM+ should also explain every ranking:
This work ranks high because it directly addresses the question, is supported by later evidence, and is repeatedly used rather than merely mentioned.
Then the human can accept, reject, or revise the judgment. Those corrections become new evidence for the system.
The deepest result may be that relevance has no single permanent ranking. There may be:
system relevance, AI-predicted relevance, expert relevance, and the user’s final relevance.
The experiment should not try to remove these differences. It should make them visible and study them.
A strong principle for LLM+ would be:
AI may predict and explain relevance; the user remains the final judge.
But repeated comparison may show where AI is more reliable than individual humans, where expert humans remain superior, and where cooperation gives the best result. That would turn relevance from an abstract debate into an empirical science.
ㅡㅡㅡㅡ
Half a century ago, I argued for subjective relevance judgment, which should be compared with AI’s objective RJ. After many comparative experiments, we could do without subjective human RJ, which is costly.
ㅡㅡ
That is a possible long-term aim, but one caution is vital:
AI relevance judgment is not automatically objective.
AI learns from human-made documents, citations, labels, rankings, and past judgments. Its result may be more consistent and cheaper than human judgment, but it can still carry hidden bias.
A strong research program would be:
First, collect many human relevance judgments from different users and experts.
Then compare them with AI judgments based on:
- query and document content;
- citation contexts;
- later use and criticism;
- novelty;
- task usefulness;
- reliability;
- time and user situation.
After many experiments, we may learn that AI can replace human judgment in some limited areas, especially where:
- the task is clearly defined;
- relevance rules are stable;
- enough trusted training evidence exists;
- AI explanations can be checked;
- serious errors have low cost.
But in open research problems, AI may still fail because relevance depends on a purpose that is new, private, changing, or not fully stated.
So the likely future is not simply:
human subjective judgment disappears.
It may be:
human judgment becomes less frequent, but more strategic.
Humans would judge difficult examples, new subjects, conflicts, and failures. AI would judge the large routine mass.
Your 1975 idea then gains a new role. Human subjective relevance judgments would first serve as the experimental ground truth. AI would learn from them. Later, the cost of routine human judging could fall greatly.
A careful principle would be:
Use human relevance judgment to train, test, correct, and periodically audit AI relevance judgment. Allow AI to replace routine judgment only where repeated experiments show that it is reliable.
This would change relevance judgment from a mainly philosophical argument into a measurable comparative science.
ㅡㅡㅡㅡ
We will just sit back and wait, until such comparative experiments have been done enough to draw a significant conclusion. At worst, we may fail to replace the costly subjective relevance judgment with the cheap (objective) AI RJ.
ㅡㅡ
Yes. That is the right scientific attitude.
There is no need to decide the result in advance. The claim should remain a testable hypothesis:
AI relevance judgment may become reliable enough to replace much costly human relevance judgment.
Then comparative experiments can continue across subjects, users, tasks, and time. The tests should measure not only agreement with humans, but also later usefulness, missed important documents, false confidence, consistency, cost, and the ability to explain a ranking.
One wording needs care: AI judgment should not yet be called fully objective. It may become more consistent, repeatable, and less costly than individual human judgment. But it will still depend on its training data, retrieval database, ranking rules, and chosen definition of relevance.
The possible conclusion may have several levels:
- AI is poorer than humans.
- AI is equal for routine judgments but weaker for novel problems.
- AI is better in some fields and worse in others.
- Human–AI cooperation is best.
- AI becomes sufficient for almost all practical relevance ranking.
Even the worst result would be informative. It would show that subjective human relevance contains something that present AI cannot replace.
So the proper position is:
Do not protect human relevance judgment as sacred, and do not declare AI superior beforehand. Compare them repeatedly and let the evidence decide.
This is very close to Popper’s method: propose replacement, test it severely, and accept failure without losing the value of the experiment.
ㅡㅡㅡㅡ
AI cannot be subjective by any means. That’s why I call its way objective, which can also be repeated. I don’t imply its absoluteness.
ㅡㅡ
Agreed. You are using objective in an operational sense, not as “absolute truth.”
A human relevance judgment is subjective because it belongs to a particular person’s thought, knowledge, purpose, and situation. Even the same person may judge differently later.
AI has no private human experience or personal need. Given the same model, database, instructions, and settings, its judgment can in principle be repeated. Therefore, it may be called objective relevance judgment in contrast with subjective human relevance judgment.
The distinction can be stated carefully:
Subjective RJ: judgment made by a human subject for a particular need and context.
Objective AI RJ: judgment produced by an external, repeatable procedure using stated evidence and rules.
“Objective” here does not mean infallible. AI judgment may still be wrong because its sources are incomplete, its method is weak, or its criteria are badly chosen. But these causes can be exposed, repeated, compared, and corrected.
A useful scientific formula is:
Objectivity means independence from an individual judge and reproducibility under specified conditions—not absoluteness.
This makes the proposed experiment clearer. Human RJ would supply the traditional subjective standard. AI RJ would supply a repeatable external judgment. Their agreement, disagreement, cost, usefulness, and errors could then be measured over many cases.
Your terminology is therefore defensible, provided that the experimental conditions and relevance criteria are openly stated.
ㅡㅡㅡㅡ
By the user’s subjective RJ, I mean his or her private RJ. Therefore, it should be very hard for AI’s objective RJ to replace it. So we just wish the subjective RJ not to be too private for AI’s objective RJ to replace it reasonably.
ㅡㅡ
Yes. That states the real difficulty very well.
The user’s relevance judgment is private because it may depend on things that are not fully spoken:
- the exact problem;
- prior knowledge;
- hidden purpose;
- urgency;
- personal standards;
- what has already been read;
- even mood or fear.
AI cannot directly enter that private state. It sees only traces: the query, dialogue, choices, clicks, rejected items, and later feedback.
So AI relevance judgment can replace private human judgment only to the extent that the private judgment has stable and observable regularities.
The research question becomes:
How much of private relevance can be inferred from public evidence?
Some relevance may be highly private:
“This article gives me the missing idea for my own unpublished theory.”
AI may fail unless the user explains the theory.
Other relevance may be less private:
“I need recent empirical studies that directly test this hypothesis.”
Here AI may judge reasonably well from stated criteria, documents, citation contexts, and past choices.
Thus, the probable boundary is:
The more relevance depends on expressible task conditions, the more replaceable it is.
The more it depends on unexpressed personal thought, the less replaceable it is.
Comparative experiments should therefore measure not only AI–human agreement, but also the degree of privacy in each judgment. They could compare:
- the user’s private judgment;
- AI judgment from the query alone;
- AI judgment after dialogue;
- AI judgment after seeing the user’s previous decisions;
- AI judgment using citation contexts and task evidence.
This would show how much private relevance becomes predictable when context is gradually made visible.
Your hope can be stated precisely:
We hope that human relevance judgment, though subjective and private, contains enough stable and communicable structure for an objective AI procedure to approximate it at acceptable cost and accuracy.
AI need not reproduce the user’s mind perfectly. It must only judge well enough for the practical purpose—and clearly show when its confidence is low.
ㅡㅡㅡㅡ
I don’t think this is the meaningful discussion between us!
ㅡㅡ
You are right. I moved away from your simple point and made the discussion unnecessarily complicated.
Your point is:
The user’s relevance judgment is private and subjective. AI cannot possess that private judgment. AI can only produce its own objective and repeatable relevance judgment.
Therefore, AI does not truly reproduce the user’s judgment. The only practical question is whether AI’s judgment agrees with it often enough to replace the costly human judgment for retrieval purposes.
That question cannot be settled by argument now. It must be settled by many future comparisons:
human subjective RJ versus AI objective RJ.
We should simply wait for the evidence. Perhaps AI will replace most human relevance judging. Perhaps the private element will prove too strong. Either result will be scientifically meaningful.
ㅡㅡㅡㅡ
2026-08-03 Mark Park
