NeoARV About

Individual vs Group Performance1

Spoiler alert: when done correctly, group forecasts outperform even the strongest members individually. That result is measured,2 replicated,3 and old enough to have a theorem behind it.4 You might be a superstar. Combined the right way, the group will still outperform you,2 and this page shows how and why.

The gut instinct says otherwise: averaging sounds like it can only drag a talented individual's prediction down toward the middle of the crowd, and the middle of an unselected crowd is chance. That instinct is right, except for one critical thing. There are many ways to compute an average, and some of them, including the ones we use, are how a group forecast comes to outperform its strongest member.

“Pooling unselected crowds dilutes talent. Pooling qualified, independent performers concentrates it.”

OK, then how do you combine them?

Every session in a forecast is worked blind and alone. Participants never see one another's sketches, notes, or calls, and each participant is assigned their own unique target image set, so the question of leakage, a viewer sharing their sketch or notes with others, has no basis for affecting anyone else's session: their sketch is for an image unrelated to anyone else's. There is no discussion to converge on and no consensus to anchor to, so the classic failure of groups, people influencing people, is absent by construction. What remains is a set of fully independent readings of the same question.

Then two filters run before any session counts toward the forecast. First, an independent judge scores every session on the SRI scale, and a session below the qualifying score does not count at all. Second, the qualifying sessions are weighted by that score: an SRI 7 session carries far more weight than an SRI 3. The figure you see is a filtered, weighted combination of independent qualified sessions, and the live breakdown shows each score tier's split separately, so a strong session is never hidden inside an average.

But we can and will do much better than this

The superior weighting is based on a viewer's track record. With a track record, we can compute the viewer's consistency: how tightly their results hold to their own overall hit rate from one stretch of sessions to the next. Two viewers can share the same lifetime rate while one delivers it steadily and the other swings between hot and cold streaks that merely average out the same, and the steady viewer's next call deserves more weight. Estimating that spread requires more history than estimating the rate itself, which is why score weighting, the strongest evidence a single day can offer, is the right method on day one. We run the best methods our data can support today, and as participation builds the record, we will upgrade to the better ones it makes possible.

Why pooling qualified viewers raises accuracy

Pooling selected calls, however, is a completely different story. The mathematics here is old and settled. Condorcet showed in 1785 that when independent judges each decide correctly more often than chance, the majority grows more accurate as the group grows, not less (jury theorems, Stanford Encyclopedia of Philosophy). The result has a modern empirical mirror: in the multi-year IARPA forecasting tournaments, group forecasts built from the best forecasters outperformed even the strongest members individually (Mellers and colleagues, Psychological Science, 2014; Tetlock and Gardner, Superforecasting, 2015). Pooling unselected crowds dilutes talent. Pooling qualified, independent performers concentrates it.

“Group forecasts built from the best forecasters outperformed even the strongest members individually.”

The mechanism is error cancellation. Skilled viewers still miss, and their misses land on different days. On any given question, most of a qualified pool is having an ordinary day, so the combined call absorbs each member's off day without inheriting it. No individual, at any skill level, can supply that correction alone.

A virtual miss-detector to boost your effective hit rate

“From the inside, a miss arrives with the same conviction as a hit.”

If you consistently hit three out of four, you carry a question that talent alone cannot answer: which of my next four sessions is likely to be the miss? From the inside, a miss arrives with the same conviction as a hit. On a day when your qualified session points one way while the other top-scoring sessions and the weighted majority point the opposite way, you are holding information about yourself that exists nowhere else: this may be your one in four, and today may be a day to stand aside. Viewers who act only on their unflagged days do not change their true hit rate. They change the hit rate of the days they act on, and that is the number a real decision rides on. Stand aside on your flagged days and your effective hit rate on the days you do act climbs toward one hundred percent. That is the virtual miss-detector, and it is the one instrument no individual, at any skill level, can build alone.

Your results remain fully yours either way. Every session shows its SRI score, your history builds your own track record, and nothing obliges you to weight the group above your own signal. Participating in a forecast simply means that while you work alone, you also get to see what every other qualified, independent session concluded about the same question.

The potential holes, and how we're ready for them

The group's power comes from a simple fact: skilled viewers miss on different days, so the group absorbs any one member's off day. That protection has one soft spot: anything that could give every viewer a bad day at once. But what could give every viewer a bad day at once? The one suspect the research keeps pointing to is geomagnetic activity. There is no scientific consensus on it, and we treat it as an open question, so we cover it from both sides: we watch the field around the clock and warn you before a session when conditions look bad (see Knowing When Not to View), and the conditions around every session go into the record, which is exactly the evidence that can settle the question for good.

The other soft spot is time: any weighting is only as strong as the history behind it. Until enough records exist, score weighting carries the load. The consistency weighting described above will go live over time as the record grows. This shift from the promising to the proven will only strengthen our forecasts.

1. Special thanks to The_Liminal_Journey, whose feedback based upon accurate observations and hard-won experience led to the creation of this page. 2. Measured across the multi-year IARPA geopolitical forecasting tournaments: Mellers et al., Psychological Science, 2014. 3. Replicated across tournament years, with the top forecasters' team advantage holding season after season: Mellers et al., Perspectives on Psychological Science, 2015; Tetlock and Gardner, Superforecasting, 2015. 4. Condorcet's jury theorem, 1785: Jury Theorems, Stanford Encyclopedia of Philosophy.

Connection interrupted

Reconnecting

Still trying to reconnect

Reload required

Attempt of

Trying again in seconds

Refresh the page to restore your session.