The easiest thing to do with 12 forecasts is average them. It is also how you erase the most useful information.
We built the first version of our market research system because no analyst can read everything. Earnings transcripts, filings, macro releases, industry reports, price history, technical signals, management commentary, and competing valuation frameworks arrive faster than a human team can absorb them.
AI helps with scale. It does not solve judgment.
In fact, the moment a system produces one clean answer is often the moment it becomes least honest about what it knows.
Our July research snapshot used 12 structured voices across eight analyst frameworks. Some emphasized valuation. Others prioritized quality, catalysts, macro sensitivity, or downside protection. The obvious product decision would have been to show readers the average forecast and hide the arguments underneath it.
Instead, three cases convinced me that disagreement deserved its own interface.

Read only the consensus rating and these look like one moderate opportunity and two non-decisions. Read the disagreement and they become three entirely different research problems.
Case one: the average said Buy, but most voices did not
SpaceX Class A received a Buy consensus rating. Yet only two of the 12 voices were positive: one Strong Buy and one Buy. One was Neutral, six were Partially Sell, and three were Sell All.
How can a mostly negative vote produce a positive result?
Because a forecast is not an election. The magnitude of the bullish cases mattered.
The optimistic thesis saw a rare orbital launch and logistics position with the potential to compound through scale, launch cadence, satellite services, and an expanding infrastructure role. The negative voices focused on immediate cash burn, maintenance capital expenditure, valuation compression, and pressure related to liquidity and tradable supply.
The average combined those views into 10.9% modeled annualized return. That number is mathematically valid. It is also incomplete.
The disagreement tells us the result is being pulled by a small number of high-magnitude scenarios. It asks the investor to inspect the tails: what probability did the bullish frameworks assign to a much larger future, and what probability did the bearish frameworks assign to capital intensity overwhelming that future?
Without the distribution, “Buy” sounds like broad support. With it, the rating becomes a fragile balance between two incompatible worlds.
Case two: the voices agreed on direction, but not on the future
PetroChina’s pattern was different. Nine voices were positive and three were negative, yet the overall rating remained Neutral. Direction consistency was just 30.7%, and the median absolute deviation of annualized forecasts was 13.2 percentage points.
The dispute was not simply about whether oil prices would rise next quarter. It was about which business PetroChina becomes over five years.
One side emphasized electric-vehicle adoption, stranded refining capacity, political constraints, and the risk that hydrocarbons absorb capital in a declining end market. The other emphasized natural gas demand, sovereign energy security, domestic infrastructure, and the possibility that state-owned enterprises receive a valuation reset as distributions improve.
Both arguments can cite real trends. They assign different weights and different clocks to them.
This is a common failure in AI consensus systems. A model can synthesize two paragraphs perfectly and still obscure that one thesis operates over 18 months while the other operates over a decade. The conflict is not factual. It is temporal.
A useful interface should expose that difference. Otherwise, the neutral rating looks like uncertainty when it may actually represent two high-conviction views using incompatible horizons.
Case three: the expected return looked attractive because the range was enormous
NuScale Power produced the most dramatic dispersion of the three. The model estimated 18.0% annualized return and 128% compounded return over five years. The consensus rating was Neutral. Risk Pressure was 98.9, and annual forecast MAD reached 27.8 percentage points.
The bullish case is easy to understand. AI data centers may require enormous amounts of reliable, round-the-clock power. Small modular reactors could address part of that need, and regulatory progress can create a meaningful barrier to entry.
The bearish case is just as concrete. Commercialization is capital-intensive. Revenue may arrive later than financing needs. Repeated equity issuance can transfer much of the eventual business success away from today’s shareholders. A liquidity squeeze could invalidate the thesis before demand has time to prove it.
This is not a mild disagreement around a stable center. It is a dispute about survival, financing, and timing.
Calling the result Neutral is defensible. Showing only Neutral is not.
Disagreement has more than one dimension
Investment systems often treat dispersion as one statistic. In practice, I have found at least three forms of disagreement worth separating.
First is directional disagreement: do the voices expect positive or negative outcomes?
Second is magnitude disagreement: they may agree that an asset rises but disagree on whether the outcome is modest or transformative.
Third is causal disagreement: they may predict the same return for completely different reasons. That matters because new evidence can invalidate one cause without touching the other.
Researchers have studied forecast dispersion for decades. A frequently cited 2002 Journal of Finance paper by Karl Diether, Christopher Malloy, and Anna Scherbina used analyst-forecast dispersion as a proxy for differences of opinion and found a relationship with subsequent stock returns in its sample. That paper does not mean disagreement is a universal trading signal. It does establish that the distribution of opinions can contain information the mean does not.
AI multiplies this challenge. Twelve voices can cover more analytical territory, but they can also share training data, data vendors, prompts, or hidden assumptions. Apparent consensus may be correlated consensus.
The right response is not to pretend the voices are independent human analysts. It is to document provenance, preserve the individual arguments, and make common inputs visible. That is why our iPulse AI data-provenance methodology treats traceability as part of the output rather than a footnote.
What a responsible AI research interface should show
NIST’s AI Risk Management Framework emphasizes validity, reliability, transparency, explainability, and interpretability as characteristics of trustworthy AI. Those words can sound abstract until you apply them to an investment screen.
For me, they translate into a practical minimum:
1. Show the consensus, but also show every underlying rating.
2. Separate agreement on direction from agreement on magnitude.
3. Identify the driver, friction, upside catalyst, and tail risk behind each thesis.
4. Preserve source dates and distinguish facts from model inference.
5. Make it possible to ask what evidence would change the result.
The SEC has warned investors not to rely solely on AI-generated information and to verify claims using multiple sources. That advice should shape product design, not merely legal disclaimers.
An AI system should help a person investigate a disagreement. It should not use polished prose to make the disagreement disappear.
Consensus is useful when it remains reversible
There is still a place for one score. Investors need to scan, rank, and decide where to spend limited attention. The problem begins when the score cannot be unpacked.
SpaceX, PetroChina, and NuScale reached different summary ratings through very different distributions. One was driven by a small bullish tail. One was divided by competing time horizons. One carried an attractive average inside a survival-level range of outcomes.
No single statistic can communicate all three.
The future of AI research will not be won by the system that sounds most certain. It will be won by the system that helps a human locate uncertainty quickly, understand why it exists, and decide which evidence deserves attention next.
The average is useful.
The argument underneath it is where the intelligence begins.
This article is for educational and informational purposes only. Model outputs are uncertain estimates based on a July 2026 snapshot and may change. Nothing here is investment advice, a recommendation, or a guarantee of future performance.
Source Notes
- iPulse AI, Batch 6 original consensus snapshot, July 2026.
- NIST AI Risk Management Framework: https://airc.nist.gov/airmf-resources/airmf/
- SEC Investor Alert, Artificial Intelligence and Investment Fraud: https://www.investor.gov/introduction-investing/general-resources/news-alerts/alerts-bulletins/investor-alerts/artificial-intelligence-fraud
- Diether, Malloy, and Scherbina, “Differences of Opinion and the Cross Section of Stock Returns”: https://onlinelibrary.wiley.com/doi/abs/10.1111/0022-1082.00490
- iPulse AI data provenance: https://ipulseai.com/methodology/data-provenance



