How the integrity model works

Why proof beats reputation

Every marketplace eventually discovers that its rating system is measuring compliance with the rating system. The alternative is not better scoring. It is asking for evidence.

Every marketplace arrives at the same problem in roughly its second year. Somebody is gaming the system, the obvious response is a better scoring mechanism, and the better scoring mechanism gets gamed within a quarter.

The reason this keeps happening is that the industry has standardised on measuring the wrong quantity. Reputation systems measure how well a participant has satisfied the reputation system. That correlates with doing good work only for as long as nobody is deliberately optimising against it, and the moment accounts acquire resale value, somebody is.

0ratings involved in deciding whether a submission is approved
48hbefore unreviewed work approves without any human judgment
72hto appeal, decided by someone other than the rejecting party

The three defences that do not hold

Worth going through them individually, because each one is genuinely useful for something and each one is routinely asked to do a job it cannot do.

Reputation scores

Useful as a rough prior about somebody you have never worked with. Useless as a gate, because the score is a lagging summary of past interactions with the scoring system, and every incentive points at maintaining the number rather than at the work. Mature marketplaces end up with a compressed band near the top where nearly everybody sits, at which point the score distinguishes nothing.

Identity verification

Genuinely necessary and frequently mistaken for sufficient. Confirming a real person exists behind an account is a different question from whether that person did the work. Identity is a property of the account. The thing being paid for is an event.

Review and rating systems

The most gameable of the three, because a review is a claim with no attached evidence and it is cheap to produce. An entire industry exists to manufacture them, which is the clearest possible demonstration that the signal has been fully separated from the thing it was supposed to measure.

What they share

All three verify the participant rather than the work. That is the structural error. Verifying the participant scales badly and degrades over time, because participants are persistent and can be optimised. Verifying the work does not degrade, because each piece of work is verified separately and there is nothing to accumulate.

Reputation asks whether to trust this person. Proof asks whether this specific thing happened. Only the second question has an answer.

What proof-based verification requires instead

The requirement is narrow and it is deliberately narrow: an artifact produced as a by-product of doing the work, which would be difficult or pointless to produce without doing it.

That last clause is doing all the work in the definition. A worker writing "completed successfully" produces an artifact, technically. It is worth nothing, because producing it costs nothing and requires no relationship to the task. A screen recording of a signup flow on a particular device in a particular country is expensive to fabricate and cheap to produce honestly, and the gap between those two costs is the entire mechanism.

Five proof types plotted by how quickly they can be produced against how strong they are as evidence.
Five proof types plotted by how quickly they can be produced against how strong they are as evidence.

The design question for any given task is never which proof is strongest. It is which proof is strong enough while still being reasonable to ask for, because a requirement heavy enough to deter fraud and honest workers equally has solved nothing. Where each format is appropriate works through that trade-off task type by task type.

Cost, not impossibility

Anybody claiming their verification cannot be defeated is either not thinking about it or selling something.

Proof can be fabricated. Recordings can be staged, photographs can be reused, metadata can be edited. The honest framing has never been that fraud is impossible; it is that fraud can be made more expensive and slower than simply doing the work, at which point the rational fraudster does the work instead, and that is a completely acceptable outcome.

  1. Raise the cost of faking

    Instance-specific requirements, evidence captured during the task rather than assembled afterwards, details that cannot be known without having been there.

  2. Lower the cost of honesty

    If the honest path is tedious enough, people route around it, and the workarounds look identical to fraud in the logs. Friction applied evenly hurts the wrong side.

  3. Keep the gap wide

    The only number that matters is the difference between those two costs. Everything else in an integrity model is a way of widening it.

Where this breaks down is high-value, low-volume work, because a large enough payment justifies a large enough fabrication effort. The answer there is not cleverer proof requirements, it is human review of the specific submission, and pretending otherwise is how platforms end up automating their way into confident wrong decisions.

The dispute mechanism is part of the integrity model

This is the part most fraud discussions skip entirely, and it matters more than any detection method.

A verification system produces judgments, some of which will be wrong. What happens next determines whether the system is trusted, and trust is what determines whether good workers stay. A platform with excellent detection and no appeal route is a platform where being wrongly rejected is unrecoverable, and the people who notice that first are the ones with options.

The 48-hour review window and every branch it can take.
The 48-hour review window and every branch it can take.

So the mechanism runs in the other direction from what fraud prevention usually implies. Unreviewed work approves automatically after 48 hours, which means a buyer who simply stops responding cannot hold a payment. One correction request is permitted rather than unlimited rounds. And a rejection can be appealed within 72 hours to an administrator rather than to the party who rejected it.

Each of those rules costs the platform something. Automatic approval means occasionally paying for work a buyer would have disputed given more time. Independent appeal means administrator hours. They are worth it because the alternative — a marketplace where workers price in the risk of not being paid — is more expensive in a way that never shows up as a line item. How appeals are actually decided covers what evidence tends to settle them.

Refusing work is a fraud control

The least discussed integrity mechanism is the list of work a platform will not accept, and it is probably the most effective one.

We turn down fake reviews, undisclosed paid endorsement, artificial engagement, bulk account creation, and anything else requiring a person to misrepresent who they are or why they are saying something. This is a meaningful amount of revenue declined repeatedly rather than a policy written once.

The reasoning is not primarily ethical, though the ethics point the same way. It is that verification and undetectable deception are mutually exclusive products. Our whole mechanism certifies that a real person genuinely did a specific thing. Work whose entire purpose is to be indistinguishable from something it is not cannot be certified, because the deliverable is the deception itself.

A verification system that will also certify a lie has not compromised slightly. It has stopped being a verification system.

There is a practical consequence that matters more than the principle. A marketplace accepting manipulation work builds infrastructure that is good at it, recruits workers who are willing to do it, and attracts buyers who want it. Every one of those is hard to reverse. The platform that decides in year three to clean up finds that its workforce, its tooling and its customer base were all selected for the opposite.

Where this model is weak

Four places, stated plainly, because a page claiming an integrity model has no weaknesses is itself a signal.

Subjective work is hard to verify. Proof establishes that something was done, not that it was done well. Judgment about quality still sits with a human reviewer, and human reviewers are inconsistent.

Collusion between a buyer and a worker defeats most of it. If both sides want a particular outcome recorded, the platform is verifying a cooperative fiction. Volume patterns catch some of this; nothing catches all of it.

Proof requirements set too low by the buyer produce weak verification, and buyers routinely set them too low because a lighter requirement gets more people to accept the task. The platform can advise. It cannot fully override somebody's judgment about their own work.

And an appeal decided by an administrator is only as good as that administrator. Independence removes the conflict of interest; it does not confer omniscience.

None of these are solved by more sophisticated detection. Three of the four are made worse by it, because automated confidence in a wrong judgment is harder to appeal than an obviously human one.

Four questions worth asking any marketplace

Regardless of which platform you are assessing, these separate a described integrity model from an actual one.

What evidence is required, and is it set before the work is posted or negotiated after somebody has already done it? Afterwards means the bar can be raised retroactively.

Who decides a dispute, and are they the same party as the one who rejected the work?

What happens when a buyer stops responding — if payment requires them to act, the worker is carrying a risk nobody disclosed.

And what does the platform refuse to sell? Somewhere that will take anything has told you precisely what its verification is worth.

Every one of those is answerable in a sentence, and the answers describe how a marketplace actually behaves far better than a trust page does. Ours are set out in the mechanism end to end and enforced through the compliance and audit trail.

Common questions

Why do reputation scores fail?

Because they measure the wrong thing. A score reflects how well somebody has satisfied the scoring system, which correlates with doing good work only while nobody is optimising against it. Once accounts have value, the score becomes a target rather than a measurement.

Is identity verification enough?

No. Identity verification answers whether a real person exists behind an account. It says nothing about whether that person did the specific work they were paid for, which is the question that actually matters.

What is proof-based verification?

Requiring an artifact produced as a by-product of doing the work — a recording, a photograph, a file with metadata, a response only somebody in that market could give. Something that would be difficult or pointless to produce without having done the thing.

Can proof be faked?

Some of it, some of the time, and the honest framing is cost rather than impossibility. Good proof requirements make faking more expensive and slower than doing the work, which removes the incentive rather than the theoretical possibility.

Why refuse work that pays?

Because verification that will also certify a deception is worthless. A platform running both verified work and undetectable fake work has one infrastructure producing both, and nobody can tell which they received.

What should a buyer ask a marketplace about fraud?

What evidence is required, who decides a dispute, what the platform refuses to sell, and what happens when a buyer stops responding. The answers to those four questions describe the integrity model more accurately than any trust page.

Assessing whether the verification actually holds?

Ask us the hard version of the question. We would rather explain where the model is weak than have you discover it later.

Ask us directly