TheJobsMarket
Finding Out What a Job Actually Pays

Self-Reported Salary Data vs Survey Data: Which to Trust

One asks employers what they pay. The other asks workers what they earn. Both are useful, both are biased, and the biases run in opposite directions.

Short answer

An employer survey describes everyone doing the work, including people whose pay has drifted, so it tends to read low. Self-reported data describes whoever chose to answer, which skews toward recent job-changers and toward the site's own audience, so it tends to read high. Use the survey to judge whether an offer is ordinary; use self-reported data when you need something about a named employer, and weight it by how many reports sit behind the figure.

Who is in each dataset

Two sources give you two different numbers for the same job, several thousand dollars apart, and both look official enough. Deciding which to believe is not really about which site is better run. It is about who ends up inside each dataset and how they got there. Once you know that, the disagreement stops being confusing and starts being useful.

An employer survey samples establishments and asks what they pay for defined occupations, with a response obligation and a consistent classification behind it. The sample is designed rather than volunteered, and it reaches employers who have never heard of a salary website. That includes rural areas, small firms and industries nobody writes about. Everyone doing the work is in scope, including people whose pay has drifted below the market over a long tenure.

A self-reported dataset contains whoever chose to submit an entry. Nobody designed that sample, nobody can weight it properly, and the people who submit differ systematically from the occupation as a whole. That does not make the numbers false, and it does change what they describe. They describe the submitters, which is a narrower and more interesting group than it first appears.

Which direction the bias runs

The employer survey tends to read low against a live offer, because it includes people who have been in the same job for a decade. Those people are frequently paid below what a new hire would command, through no fault of their own. So a median that includes them sits under what a company would actually have to pay you today. That is not a flaw, and it does mean an offer above the median is not automatically generous.

Self-reported data leans the other way, and for more than one reason at once. People pleased with their pay submit in order to signal it. People researching a move submit as the price of access, and they skew toward the ambitious, the mobile and the urban. Meanwhile everybody in a stable job who has never visited a salary site is missing entirely, and they are a large share of most occupations.

The platform shapes the answer

A site built for software engineers reports a different figure for the same job title than a general careers site does. Both can be describing their own users honestly while disagreeing sharply with each other. Neither is lying, and neither is describing the occupation as a whole. The figure is a fact about that platform’s audience.

So the question is never simply what a site says, but who uses it. A number from a platform serving a high-paying niche is a number about that niche, and importing it into a different context inflates your expectations. Check who the site is built for before deciding what its figure means. That takes about a minute and it changes how much weight the number deserves.

What self-reported data is genuinely better at

Recency is the first advantage and it is a real one. Submissions arrive continuously while survey data describes a period about a year back. In a market that has moved recently, the self-reported figure catches something the survey cannot yet see. That is worth a great deal when you are negotiating this month rather than reasoning about a decade.

Granularity is the second. Employer name, internal level, and the components of a package are all things an official survey never publishes and never will. Total compensation is the third: equity, bonus and sign-on money sit almost entirely outside the official wage figures. For the question of what one specific company pays at one specific level, self-reported data is the only source that exists.

What survey data is better at

It describes the level and shape of an occupation across the entire country, including the parts nobody writes about. It reaches rural areas, small employers and unglamorous industries that no salary platform will ever cover. If you work outside a major metro or outside a fashionable industry, this is the source that actually contains you. That coverage is not something a larger self-reported sample can substitute for.

It also carries consistent occupational definitions, which means comparing two occupations is a meaningful exercise. Comparing two job titles across self-reported entries frequently compares two different levels that happened to be typed the same way. One company’s Analyst II is another’s Senior Analyst, and the dataset cannot tell. The classification exists precisely to remove that problem.

How to use both without averaging them

Do not split the difference between them. They answer different questions, and the midpoint of two different questions is not an answer to either one. Averaging feels balanced and produces a number that describes nothing in particular. It is the most common mistake people make with these two sources.

Use the survey for the range and where the middle sits, because that is what it measures well. Use self-reported entries for what a named employer pays and what the package contains. Where the two disagree by a wide margin, the self-reported sample is usually skewed toward one end. In that situation the survey is the better guide to what is typical, and the self-reported figure is the better guide to what is possible.

Sample size is not the reassurance it looks like

A self-reported figure built on eight submissions is close to meaningless, and most sites do not show the count. That is the first thing to look for and the thing most often missing. Where a count is shown, treat anything under a couple of dozen as an anecdote rather than a benchmark. It may still be the only information available, and it should not carry more weight than that.

A large count does not fix the underlying problem either, which is where the real misunderstanding sits. Ten thousand self-selected entries are still ten thousand people who chose to submit, and the selection runs through all of them equally. Sample size cures noise and does nothing at all for bias. Where the count is hidden, assume it is small, because sites display the number when it flatters them.

The practical rule

Before an interview, know the survey median and the full percentile spread for your occupation in your metro. That is your anchor, it is checkable, and it is difficult for anybody to wave away. Then use self-reported data for the specific employer, which tells you where in that range to aim. The two together give you a defensible floor and an informed target.

The order matters more than it sounds. Bringing an official figure into a negotiation is hard to dismiss, because the person opposite can look it up and find exactly what you said. Bringing a screenshot from a salary site invites a conversation about the source instead of about your pay. Lead with the number that survives being checked, and keep the other one for deciding what to ask for.

Common questions

Why is Glassdoor-style data usually higher than government data?

Because people submit salaries when they are testing the market, and recent movers are paid more than long-tenured incumbents in the same role. The survey includes everyone; the self-report over-represents people who just changed jobs.

How many reports do I need before a self-reported figure is useful?

There is no magic threshold, but a figure backed by single digits is an anecdote. Look for whether the site tells you the count at all — many do not, and that is itself informative.

Can I average the two sources?

No. They describe different populations, so the midpoint describes nobody. Work out which population you belong to and use the source built to describe it.

Is self-reported data verified?

Almost never. Some sites ask for a document upload for a subset of submissions, but the bulk of what you see is unverified and self-selected.

What does the gap between the two tell me?

Roughly the premium that changing employer commands over staying. If it is wide in your occupation, that is a strong argument for testing the market rather than asking for a raise.

Is self-reported salary data reliable?

For a specific employer and package, it is often the only source. For what an occupation typically pays, it overstates — the people who submit are not a designed sample.

Should I average survey and self-reported figures?

No. They answer different questions, and the midpoint of two different questions answers neither.

How many submissions make a figure meaningful?

Treat anything under a couple of dozen as an anecdote. A large count still does not fix selection bias, and sites show the number only when it flatters them.

CS

Charles Slocs

Data and research

Charles Slocs builds the data side of this site — pulling the federal wage and employment series, matching job titles to occupation codes, and working out what the numbers do and do not support. He writes the pages that are mostly a question about evidence: what a survey measured, how wide the spread really is, and which published figure is out of date.

All articles by Charles Slocs →