Modification to the Sapienza probability adjustment for trust to include lying and bias
More actions
Main article: Trust
Review of Sapienza's Probability Adjustment Equation
Sapienza’s paper (https://ceur-ws.org/Vol-1664/w9.pdf), which is the basis for our modeling approach, uses the notion of trust between nodes to adjust the node’s probability of getting a particular answer. Each node is first assigned probabilities for a particular outcome (eg 60% Red, 30% Blue, 10% Green). Then, based on trust, the nodes’ probabilities are adjusted (or smoothed, as Sapienza puts it) up or down based on the following eqn:
where
is the adjusted probability
is the nominal probability ()
is number of choices
is Trust
is the raw probability, before adjustment
This equation, it should be noted, looks a little different than the way Sapienza presents it because he uses a probability distribution and we use discrete probabilities. In any event, this equation is then used in the Bayes equation to calculate the combined probability, given several sources.
The Sapienza adjustment for probability was shown in a previous post to be equivalent to having random answers for the “untrustworthy” part of the trust, ie .
In other words, if I trust my friend 80% then I believe 80% of what he says represents the truth (as he sees it) and the other 20% is random. If I ask him what tomorrow’s weather will be like, rainy, sunny, or cloudy and his view is 70% cloudy, that number will be lower because he will only report that view 80% of the time. For the remainder, he will report rainy, sunny, or cloudy at random. This leads to the adjustment above:
So our friend’s 70% view is actually 62.67% due to trust. This makes sense and reduces probabilities to when Trust is zero and increases them to when Trust is 100%.
Implicit in this equation, however, is what we are doing with the “untrustworthy” part (). Sapienza, quite reasonably, chose to make this part random. His purpose was not to model these details but rather to produce a practical Bayesian model that takes into account trust at a high level.
An equation with randomness, lying, and bias
Here we extend this equation by including, along with randomness, lying and bias. A node may lie to us or produce biased answers and we’d like to take this into account in constructing the model. Bear in mind that we must do something with the “untrustworthy” part of trust. Unless trust is always 100% we must decide what represents. Sapienza chose random. Here we include lying and bias.
First let’s present the equation:
where
is modified probability
is raw probability (unmodified)
is Trust that node is reporting the truth as they see it
is extent to which is random (ie in Sapienza)
is extent to which is lying
is extent to which is biased toward the answer being calculated (ie if we are calculating the of the Red choice then is the bias toward Red – that is, the node’s answer is always Red for this portion)
is number of choices
A snippet to calculate this equation is here: https://gitlab.syncad.com/peerverity/trust-model-playground/-/snippets/137
The next sections deal with justification, derivation of the equation, how it reduces to Sapienza’s original equation when lying and bias do not exist, and some notes on how it might be further refined.
Justification
A common objection is "why would I include anyone in my network who lies to me?". The answer is that:
- Even trustworthy people lie sometimes, depending on the subject.
- It’s better to have a general model than one that implicitly assumes a behavior (in this case, random behavior).
- The model doesn’t force anyone to include lying in their trust equation. They can set it to zero ().
Bias is even more common and easier to accept in the context of trust. Almost everyone we know is biased in one way or another.
By constructing a general equation and, more importantly, showing how it is derived, we can model any scenario for .
Derivation
We will derive the general equation above using an example with numbers. Let's suppose we have 3 choices, Red, Blue, Green with the following probabilities:
We assume a trust of 70%, leaving 30% for the “1-T” portion. Let’s say 8% is random, 15% is a lie, and 7% is biased toward Blue. The lie is anything other than the truth, evenly split between all the untrue options, and bias means the node chooses Blue no matter what for that 7%. There is no bias toward the Red or Green although in theory there could be.
Note that the total of the trusts is 1, as you would expect, and that
This is why we say we are modeling the portion of the trust.
We can construct the following table based on this scenario assuming there are 120 samples, just to have some numbers to work with:
Now if we simply add up the number of reported Red, Blue, and Green results, we can write the following table:
Each term in the addition above can be expressed symbolically. The modified probability is this addition divided by the number of samples. For the reported Blue (rB) case, for instance, we obtain:
The N cancels and we can collect together the terms involving ( in denominator), ( in denominator), and (no denominator):
We note that and that :
The equation is the same whether we consider the Red, Blue, or Green outcome. In general,
Equivalence with Sapienza Trust Equation
If we remove the terms containing $T_l$ and $T_b$ we can do some algebra and return to Sapienza's original equation for modification based on trust:
We note first that :
We further note that :
Then using some algebra:
which is Sapienza’s original equation.
Possible Refinement
One way to make the new equation more realistic is to make the lying correspond to the bias. As it stands, this equation assumes that lies are evenly split between all options that are not the truth. However, if the respondent has a bias then the lie might well simply correspond to the bias, as long as the bias is not the truth.
Some observations that folks have been made about the model include:
- There will rarely be “random liars” (liars modeled by ). In practice, it seems like most liars will lie towards their bias, as you suggest above.
- A bias towards a particular choice or even just truth/false is problematic to model usefully in practice. Let’s take a simple example. Say you think a person is biased towards an authoritarian party. You would need to inspect each individual predicate related to the topic to see if the bias should be applied. Even two predicates about the same politician would need to be individually evaluated, because one predicate could be positive towards the candidate and another could be negative. So to effectively account for this kind of bias, there’s going to need to be either AI or some sort of crowdsourcing to determine which answer the bias counts towards.

