An employer survey describes everyone doing the work, including people whose pay has drifted, so it tends to read low. Self-reported data describes whoever chose to answer, which skews toward recent job-changers and toward the site's own audience, so it tends to read high. Use the survey to judge whether an offer is ordinary; use self-reported data when you need something about a named employer, and weight it by how many reports sit behind the figure.
Who is in each dataset
An employer survey samples establishments and asks what they pay for defined occupations, with a response obligation and a consistent classification. The sample is designed rather than self-selected.
A self-reported dataset contains whoever chose to submit. Nobody designed that sample, nobody can weight it properly, and the people who submit are systematically different from the occupation as a whole.
Which direction the bias runs
Both ways, which is why it is hard to correct. People pleased with their pay submit to signal it. People researching a move submit as the price of access, and they skew toward ambitious, mobile and urban.
Meanwhile people in stable jobs who have never looked at a salary site are absent entirely, and they are a large share of most occupations. The result is usually an overstatement, especially at the top of the range.
The platform shapes the answer
A site built for software engineers reports different numbers for a job title than a general careers site, because different people use them. Both can be describing their own users honestly while disagreeing sharply.
So the question is never just “what does this site say” but “who uses this site”. A figure from a platform aimed at a high-paying niche is a figure about that niche.
What self-reported data is genuinely better at
Recency. Submissions arrive continuously; survey data describes a period a year back.
Granularity. Employer name, level, and the components of a package — none of which an official survey publishes.
Total compensation. Equity, bonus and sign-on, which the official wage data largely excludes.
For “what does this specific company pay at this level”, self-reported data is the only source there is, and it is genuinely useful for that.
What survey data is better at
The level and shape of an occupation across the whole country, including the parts of it that nobody writes about. It covers rural areas, small employers, and unglamorous industries that no salary platform reaches.
It also has consistent occupational definitions, so comparing two occupations means something. Comparing two job titles across self-reported entries frequently compares two different levels that happened to be typed the same way.
How to use both without averaging them
Do not split the difference. They answer different questions, and the midpoint of two different questions is not an answer to either.
Use the survey for the range and where the middle sits. Use self-reported entries for what a specific employer pays and what the package contains. Where they disagree by a lot, that usually means the self-reported sample is skewed toward one end, and the survey is the better guide to what is typical.
Sample size is not the reassurance it looks like
A self-reported figure built on eight submissions is nearly meaningless, and sites rarely show the count. But a large count does not fix selection bias either — ten thousand self-selected entries are still ten thousand people who chose to submit.
Where a count is shown, treat anything under a couple of dozen as an anecdote. Where it is not shown, assume it is small, because sites display the number when it flatters them.
The practical rule
Before an interview, know the survey median and spread for your occupation in your metro. That is your anchor and it is defensible. Then use self-reported data for the specific employer, which is what tells you where in the range to aim.
Bringing an official figure into a negotiation is difficult to dismiss. Bringing a screenshot from a salary site invites a conversation about the source rather than about your pay.
Common questions
Why is Glassdoor-style data usually higher than government data?
Because people submit salaries when they are testing the market, and recent movers are paid more than long-tenured incumbents in the same role. The survey includes everyone; the self-report over-represents people who just changed jobs.
How many reports do I need before a self-reported figure is useful?
There is no magic threshold, but a figure backed by single digits is an anecdote. Look for whether the site tells you the count at all — many do not, and that is itself informative.
Can I average the two sources?
No. They describe different populations, so the midpoint describes nobody. Work out which population you belong to and use the source built to describe it.
Is self-reported data verified?
Almost never. Some sites ask for a document upload for a subset of submissions, but the bulk of what you see is unverified and self-selected.
What does the gap between the two tell me?
Roughly the premium that changing employer commands over staying. If it is wide in your occupation, that is a strong argument for testing the market rather than asking for a raise.
Is self-reported salary data reliable?
For a specific employer and package, it is often the only source. For what an occupation typically pays, it overstates — the people who submit are not a designed sample.
Should I average survey and self-reported figures?
No. They answer different questions, and the midpoint of two different questions answers neither.
How many submissions make a figure meaningful?
Treat anything under a couple of dozen as an anecdote. A large count still does not fix selection bias, and sites show the number only when it flatters them.