representAI

representAI / Methods

Two ways to complete missing choices

One follows the logic of your own answers. The other learns which users and pairs tend to agree. They are separate models, with separate results.

Choices across users

Use a matrix with users as rows and alternative pairs as columns. Similar users and related pairs help fill gaps.

See the pair matrix ↓

One workflow, independent experiments

Choose a study: policies, everyday annoyances, flavors, or another published context. Each study defines its own translated items and question. Country labels describe the study material, not the participant. Sessions and training cohorts are separated by study and configuration version. Test pairs never appear in that session’s learning round.

“Both” and “Neither” follow the study’s question: support for policies, annoyance for peeves, or enjoyment of flavors. Joint approval below means this joint assessment, not universally liking the items. Each middle answer records zero directional preference. The site asks no demographic questions, names, contact details, location or free text.

First: missing is not zero

An unanswered comparison has no response record. We store each actual answer by name, then derive two separate signals. This keeps “prefer B” different from “Neither,” and “prefer A” different from “Both.”

Encoding for a fixed pair A–B
AnswerDirectionJoint approval
Prefer A+1Unknown
Prefer B−1Unknown
Equal00
Both0+1
Neither0−1
No answerMissingMissing

Pair order is fixed by proposal ID, independently of left/right screen position. Equal is an observed neutral answer. A preference for one alternative does not tell us whether the person approves of both alternatives jointly.

1. Personal transitivity

Think of your learning answers as arrows: an arrow from A to B says that you prefer A. Following A → B → C allows the model to infer A → C.

ABC
AABBCC—1?0—1?0—

Read a cell as “row preferred to column.” A recorded 0 is the reverse of a known preference, not missing data. Question marks are unknown. Amber cells are implications of the chain.

One person, three alternatives. No other user’s answers are involved. Unconnected pairs stay unknown; contradictory chains are not treated as reliable predictions.

The model follows strict preference arrows and treats the three middle answers as ties in direction. Both and Neither still retain their different joint-approval meanings. Missing connections remain unknown. If a relevant chain contradicts itself, the model abstains instead of forcing a ranking.

Personal model: mathematical definition

The relation is represented by alternatives × alternatives, for this person only. Strict preference adds a directed edge; each middle answer adds a tie link. A known reverse preference can be displayed as 0, with a separate observation mask. Unanswered cells are never filled with 0.

A≻B ∧ B≻C⟹A≻CA\succ B\ \land\ B\succ C\quad\Longrightarrow\quad A\succ C

We compute reachability and whether each path contains a strict edge. A path containing a strict edge implies strict preference; paths using only Equal links can predict Equal. Other tie-only paths remain unavailable because transitivity does not determine their joint approval. If a strict cycle lies on a proposed proof, the answer is unavailable. Incomparable alternatives need not be forced into a total order.

This model can infer a preferred proposal or Equal. Transitivity alone cannot infer joint approval, so it does not predict Both or Neither. A stored confidence of 1 means a logical implication under these assumptions, not a measured probability.

2. Across-user pair relatedness

Here the columns are A–B, A–C, A–D, and so on. We maintain one sparse matrix for preference direction and another for joint approval. Rows are anonymous sessions: repeat runs are separate rows, not identifiable people.

Preference direction. Rows: u1, u2, you. Columns: A–B, A–C, A–D, B–C. Hollow question marks are missing, filled zero is observed neutral.Preference directionA–BA–CA–DB–Cu1u2you11-1011-1011??

Each column is a pair, not a single proposal. You agree with u1 and u2 on A–B and A–C. Your A–D answer is missing. The filled zeros for B–C are recorded ties.

Illustrative data. Joint approval uses a separate matrix and the same procedure. Observed values are never replaced with fabricated zero answers.

For each missing cell, one estimate comes from similar users who answered that pair. A second comes from related pairs the current user answered. We average the available estimates. Negative similarity supports an opposite answer; no usable evidence means no prediction.

This is inspired by relatedness density in economic complexity: nearby observed activity helps estimate a missing entry. Here we use masked, signed similarities for preference data, rather than the original binary co-export proximity measure. See the Atlas of Economic Complexity glossary for the original relatedness context.

Across-user model: matrices and equations

For either signal, let XX have mm sessions and nn canonical pairs. Its known values belong to {−1,0,+1}\{-1,0,+1\}. Let MM record which cells are observed.

Mup={1,Xup is observed,0,otherwise.M_{up}=\begin{cases}1,&X_{up}\text{ is observed},\\0,&\text{otherwise}.\end{cases}

Similarity uses only jointly observed cells, including observed zeros. For shared index set II, signed cosine is shrunk by overlap. Fewer than two shared observations gives no link. Two all-zero observed profiles agree; if just one has zero norm, the link is zero.

sim⁡(x,y)=∣I∣∣I∣+2 ∑i∈Ixiyi∑i∈Ixi2∑i∈Iyi2\operatorname{sim}(x,y)=\frac{|I|}{|I|+2}\,\frac{\sum_{i\in I}x_i y_i}{\sqrt{\sum_{i\in I}x_i^2}\sqrt{\sum_{i\in I}y_i^2}}
S∈Rm×m(users),T∈Rn×n(pairs)S\in\mathbb R^{m\times m}\quad\text{(users)},\qquad T\in\mathbb R^{n\times n}\quad\text{(pairs)}

Self-links are zero. Pair similarities use prior users only. Similarity between the current user and a peer excludes the target pair. We calculate only the rows of these square matrices needed for each prediction.

x^upusers=∑v≠uSuvMvpXvp∑v≠u∣Suv∣Mvp,x^uppairs=∑q≠pTpqMuqXuq∑q≠p∣Tpq∣Muq.\begin{aligned}\widehat x^{\mathrm{users}}_{up}&=\frac{\sum_{v\ne u}S_{uv}M_{vp}X_{vp}}{\sum_{v\ne u}|S_{uv}|M_{vp}},\\\widehat x^{\mathrm{pairs}}_{up}&=\frac{\sum_{q\ne p}T_{pq}M_{uq}X_{uq}}{\sum_{q\ne p}|T_{pq}|M_{uq}}.\end{aligned}

The sums include observed cells only. These are normalized matrix products; the masks prevent missing values from contributing to either numerator or denominator. An observed zero still contributes to the denominator. A zero denominator gives an unavailable estimate, never an imputed zero. The available user and pair estimates receive equal weight.

Let dd be the resulting direction score and jj the joint-approval score. Both are bounded by −1 and +1. Apply the following rules in order:

answer⁡={unavailable,d missing,A,d>0.25,B,d<−0.25,unavailable,j missing,Both,j>0.25,Neither,j<−0.25,Equal,otherwise.\operatorname{answer}=\begin{cases}\text{unavailable},&d\text{ missing},\\A,&d>0.25,\\B,&d<-0.25,\\\text{unavailable},&j\text{ missing},\\\text{Both},&j>0.25,\\\text{Neither},&j<-0.25,\\\text{Equal},&\text{otherwise}.\end{cases}

The threshold 0.25 is a modeling convention, not a calibrated probability. A directional tie without joint-approval evidence remains unavailable. Confidence is a heuristic score strength, not a measured likelihood of correctness.

What data can the models use?

Both models use the current session’s learning answers only. The personal model requires no other users. The across-user model additionally uses learning answers from at most 500 unexpired sessions of the same study version, completed before the current session began. It has no fixed minimum number of users, but each similarity needs two shared observations. Insufficient overlap produces unavailable predictions.

The current prediction-round answers never train either model. Predictions are saved before showing each new pair. Model versions, training digests, cohort cutoffs and support diagnostics are recorded for auditing. Deletion can change evidence available for later predictions; already saved predictions are immutable.

How results are scored

Accuracy requires an exact match to one of the five recorded answers: A, B, Equal, Both or Neither. Availability is shown separately. Random and population baselines remain separate: they predict one proposal and cannot predict the three middle answers. These different coverage and answer capabilities should be considered when comparing scores.

Open data, repeat runs and earlier versions

The console lists current questions and downloads completed sessions that consented to public sharing. Raw answer labels preserve Both, Neither and Equal separately; missing responses are never encoded as zero. Public files omit demographics, exact timestamps, credentials and private training provenance.

Each repeat run creates a new anonymous session. Save its predecessor’s deletion code first. Deleting a run removes future exports but cannot recall downloaded copies. Session counts are not counts of unique people, and de-identification does not guarantee that preference patterns are unrecognizable.

Political study version 4 and the new demonstration studies use the common anonymous flow. Version 3 introduced Both and the two methods above. Earlier snapshots retain their original answer options and predictors, including the version-2 low-rank session-by-proposal model. Those older results are not relabeled or recalculated. No LLM or external AI service is enabled.