I Gave Two AI Judges the Same 14 Reports. One Saw +5%. One Saw −10%.
The source evidence stayed fixed. Two forecasts reversed direction, and Bitcoin’s unanimous panel concealed a 29-point gap.Five-year projected price returns under two AI judges, with the same fourteen reports per company. Archived blinded baselines, reproduced October 3, 2026.I gave two AI judges…
Annons
Annons
The source evidence stayed fixed. Two forecasts reversed direction, and Bitcoin’s unanimous panel concealed a 29-point gap.Five-year projected price returns under two AI judges, with the same fourteen reports per company. Archived blinded baselines, reproduced October 3, 2026.I gave two AI judges the same fourteen reports about Moderna. One produced a projected five-year gain of 4.98%. The other produced a projected loss of 9.98%. When I compared Nebius, the forecast crossed zero in the opposite direction: −0.42% became +8.10%.No source report needed to change. The model deciding how much each report counted was enough. Building an agent research platform had made me pay attention to the analysts; this experiment made me pay closer attention to their chair.These results come from my archived study, When Models Judge Models. It examines 98 source reports and 56 consensus outputs across seven deliberately selected assets. The sources were generated on September 20, 2026. I reproduced the numerical results offline on October 3, with all 392 recorded checks passing.The final model looks easy to overlook because it arrives after the expensive research. The agent team has collected its perspectives; the remaining task appears to be a summary. These results put the surprise at that final step. The summarizer is choosing whose view receives authority, and that choice can reverse the number a reader takes away.The chair of the meeting controls the answerImagine fourteen analysts submitting reports to an investment committee. Some emphasize business quality, others macroeconomic forces, technological change, or hidden risks. They can all read the same company and still disagree about which forces matter most. The committee’s chair then decides how much influence each report should receive.The AI system here has a comparable structure. Its final model allocates scores to the fourteen reports. Code normalizes those scores into weights. Those weights determine the contribution of each report’s quarterly predictions to the resulting forecast path. The judge is therefore choosing influence, not merely shortening a meeting transcript.That distinction matters for any agent workflow in which a final model resolves competing recommendations. The decision can be hidden inside a persuasive synthesis: one report is described as insightful, another as too speculative, and a third as missing the main issue. Unless the influence is recorded, it is difficult to see what those descriptions changed.In this study, the influence is numerical, so we can trace it. A model’s allocation becomes a weighted quarter-by-quarter forecast and then a compounded five-year return. That makes the judge’s consequences visible in a way that a prose-only committee summary would not.That makes the chair part of the reasoning system we need to evaluate. Collecting diverse reports solves only the problem of having perspectives available. It leaves open how much power each receives, which conflicts survive the synthesis and which disappear inside its polished conclusion.Fourteen voices can share the same blind spotThe source composition needs to be understood before interpreting the results. Twelve reports use the archived Gemini family configuration. One uses the archived Astra configuration and one the archived Opus configuration. They span nine analytical frameworks and selected research modes. This is diversity of perspective, with substantial model-family overlap.The framework changes the questions asked: ownership economics, debt cycles, disruption, political incentives, and other lenses. The mode changes how evidence is acquired. A researcher can search; a thinker reasons from supplied dated context. The model is the implementation doing that work. Those are three distinct choices.That means a fourteen-report consensus is not a vote among fourteen statistically independent models. A majority can reflect shared implementation or assumptions as well as the evidence. The study does not measure those dependencies well enough to turn the report count into a probability of being right.The platform is iPulse AI, an investment research system I develop. Its value for this experiment is that the reports, influence allocations, and forecast paths can be inspected together. The experiment is about what that inspection reveals, including limits, rather than a claim that a larger team automatically produces better advice.More providers can change the available perspectives. They do not automatically turn a collection into independent evidence. If twelve reports share an underlying family, their different frameworks may still inherit related limitations. A larger panel can look broad while a common error travels through many of its seats.Agreement has a direction, a magnitude and a priceThree selected equities — Visa, SAP and IBM — have fourteen positive five-year source forecasts each. Three other equities show directional disagreement: Nebius and Moderna split seven positive against seven negative reports, while Kratos splits eight against six. Bitcoin is a separate reference case: all fourteen forecasts are positive, but their magnitudes differ widely.These categories were deliberately selected to make different kinds of disagreement visible. They are not a random sample of the market, and the complete original ranking universe is not independently reproduced in the public release. The cohort comparison should therefore be read as a contrast among these cases, not a population estimate.Five-year projected price returns under Gemini and Astra across seven selected assets. Source: archived baseline comparisons.Within the three aligned equities, the mean absolute difference between the judges’ five-year forecasts is 3.02 percentage points. Within the three divergent equities, it is 11.64 points. That is consistent with the idea that conflicting source paths give the judge more consequential choices to make.But the relationship is not a universal law. IBM’s two forecasts differ by 7.53 points despite unanimous positive source directions. Bitcoin’s gap is larger than any equity’s. Saying that reports agree about the sign tells you much less than saying they agree about the path and magnitude.A display saying “the agents agree” compresses all of this into a comforting phrase. Agreement that a return is positive leaves open how positive, by when, and under which assumptions. The next case shows how much room can remain inside unanimity.Bitcoin gets 14 votes for “up.” The judges stay 29 points apart.Bitcoin’s two blinded baseline judges project five-year cumulative price returns of 153.15% and 123.70%. Both are positive, but the gap is 29.45 percentage points. The middle half of the fourteen source terminal returns spans 64.33 points. Unanimity about direction leaves plenty of disagreement about the size of the future.The allocations help explain why a single bullish label would hide so much. Astra assigns 55% of its baseline weight to its own family’s source report. Gemini assigns about 12.86% to that same report. The two consensus paths reflect different allocations over the same source set.The weight difference is visible: one judge concentrates far more influence in the single Astra-family report. It helps describe how the two syntheses differ, though it does not by itself isolate the cause of every percentage point in the final gap. The important discovery is that unanimous direction still leaves substantial authority to the judge.Nor does it establish that either projected gain will occur. These are price paths anchored to a saved September 2026 reference, not a current trading recommendation. The forecast endpoints lie years ahead. A dramatic numerical difference is an audit finding about system behavior, not a shortcut to predictive validation.The reader needs to see both the vote and the distance. All fourteen reports projecting gains tell one story. Their spread and the judges’ different allocations tell another. Hiding the second story makes consensus sound more complete than the source evidence warrants.Visa’s quiet headline hides a rearranged committeeVisa gives the opposite lesson. Its two baseline weight allocations differ by roughly 24.4% in total variation, yet the projected five-year returns differ by only 0.24 percentage points. A considerable share of influence moves around while the headline stays nearly stationary.Total variation here means half the sum of absolute differences between normalized weights. It measures the amount of influence that must be reassigned to turn one allocation into the other. It does not measure forecasting error, and it does not describe how much a real portfolio changed.Visa supplies the reverse surprise. Large movements in influence can produce a tiny change in the endpoint when the underlying projections are similar. A nearly unchanged answer therefore cannot establish that the judge relied on the same evidence. The apparent calm belongs to the headline, rather than necessarily to the decision beneath it.That matters for the next update. A judge leaning on one report today may respond differently when that report changes tomorrow. Keeping the influence record gives us a way to investigate the reaction. Keeping only the endpoint leaves the committee’s rearrangement invisible.Ask the unchanged question againThe archive also reruns an unchanged-order blinded presentation and reverses the report order. Across the six equities, Astra’s mean five-year movement on the same-order repeat is 0.93 percentage points, compared with 2.17 for Gemini. That looks encouraging if repeatability is your main concern.Under reversed order, the comparison changes: Astra’s mean five-year movement is 2.04 points, compared with 1.37 for Gemini. A deployed system can be more repeatable under one change and less stable under another. A single badge saying “more robust” would hide that distinction.These are small descriptive comparisons. There are only two blinded realizations per asset and synthesizer, one reversal and one disclosure run. We cannot use them to estimate a precise distribution of outcomes or make a general vendor ranking. Provider tools, reasoning controls, and timing also differ.Bitcoin adds an important boundary. Its reversal and disclosure requests contain an extra checklist relative to baseline and repeat. Those contrasts are confounded. The cleaner presentation comparisons use the six equities; Bitcoin remains valid for the baseline judge comparison. Keeping the awkward exception makes the analysis stronger, not weaker.The useful next experiment would repeat the same controlled conditions often enough to estimate their variability. Here, the small paired archive supplies concrete cases and measured contrasts. It points toward the audit we need; it does not establish a universal provider ranking. The surprise is worth retaining together with that boundary.What the final answer owes the readerFirst, I want to see the source disagreement before the judge resolves it. A consensus interface should make clear whether the sources differ over direction, magnitude, timing, or underlying assumptions. The final answer needs that context if readers are to understand how much judgment was required to produce it.Second, I want the influence allocation retained. In a numerical workflow, the weights should be inspectable. In a prose workflow, the system should preserve which source claims were accepted, rejected, or used to resolve a conflict. That is a proposed design principle, not a claim that prose attribution alone guarantees faithful reasoning.Third, I want a simple comparator. What happens under equal-report weighting, and what changes when the judge changes? A comparator is not automatically superior. It gives us a way to tell whether the extra judgment produces a meaningful difference and eventually whether that difference helps.Finally, I want an outcome-evaluation plan whose rules are fixed before results arrive. The current study cannot tell us which forecast will prove better. It can tell us what must be preserved so that a later evaluation answers a stable question rather than a revised one.The breakthrough may begin with the team. The answer still passes through its chair. Moderna’s gain became a loss, Nebius’s loss became a gain, Bitcoin’s unanimous vote concealed a large magnitude gap, and Visa’s stable endpoint concealed changing influence. Those are four reasons to make the last model as inspectable as the first fourteen reports.ReferencesWhen Models Judge Models: the archived case study, dataset and limitations supporting the numerical findings.Versioned research repository: source reports, requests, allocations and offline reproduction.LLM-Blender: prior work on combining language-model outputs through ranking and fusion.iPulse AI: the research platform studied; I am the founder.Mixture-of-Agents Enhances Large Language Model Capabilities: prior work on layered agent collaboration; its results do not validate this investment system.Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena: prior evaluation research on language models acting as judges.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!I Gave Two AI Judges the Same 14 Reports. One Saw +5%. One Saw −10%. was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.