All posts

We Ran One Resume Through Three Tailoring Tools. One Raised Its Score For A Skill It Invented

One controlled resume, one job posting, three tailoring tools. One marked real skills missing, then raised its score for a skill the resume never contained.

Data 27 Aug 2026 13 min read

The final screen read 15 percent, and under it sat a list of the skills the job wanted. Exactly one of them was rendered as present: Apache Flink. The resume we had uploaded does not contain the word Flink. It never contained it, not once, in either of its two pages. The tool had offered to write a bullet, we had let it write one about Flink, we had saved that bullet, and the score went up. Kafka, Spark and Iceberg, all three of which the resume genuinely does contain, were still marked missing on the same screen.

That screen is the reason this post exists. It is one observation from a small, deliberately boring benchmark: one control resume, one job posting, four resume tailoring products, run on 27 August 2026, with every count taken from text extraction of the source document rather than read off a screenshot. The protocol, the raw numbers and the observations are published as a data file so the run can be repeated, checked, and corrected.

This benchmark is published by JobShifu, which is also one of the products in it: JobShifu's own run is labeled, and every number here traces to the published data file and the screenshots in the post [1]. Nothing below depends on trusting the narration. If a number looks wrong, the resume is a fixture, the postings are public, the tools are all free to sign up for, and the file lists what was measured and when.

How the benchmark ran

How the benchmark ran: one control resume and one job posting into Teal, Jobright, Simplify and JobShifu, with every claim checked against the source document rather than read off a screen.
The protocol in one picture.

The control resume is a fixture, not a real person: two pages, a fictional identity, a senior data platform engineer with a plausible history. Using a fixture matters more than it sounds. A real resume drifts, gets edited between runs, and makes term counts unfalsifiable. A fixture can be counted, and the counts can be published.

Those counts are the spine of the whole exercise. By text extraction of the source file, re-verified on 27 August 2026, the resume contains Kafka twice, Spark five times, Iceberg three times, Trino once and EMR twice. It contains the word Apache zero times, and the word Flink zero times. Everything a tool says about this resume can therefore be graded, because the ground truth is a number rather than an impression.

The job posting is the other fixed input. For the three third party tools it is a ServiceNow role in San Diego, California, titled Senior Staff Data Platform Engineer, whose title goes on to name Kafka, Apache Iceberg and Apache Spark. It was chosen precisely because the overlap with the control resume is high and obvious. A human reading the two documents side by side would call this a strong match on the technical core, and would notice immediately that the candidate has never touched Flink.

The rule for every claim in this post: if it is a count, it comes from extracting text out of the source document. If it is a tool's verdict, it comes from a screenshot with a date on it. The two are reported separately and never averaged into a single judgment. That separation is the only reason the finding in the lede is visible at all.

Teal reads the posting literally

Teal's match score screen before any edits: 9 percent, Hard Skills 3 of 30, with Apache Kafka, Apache Flink, Apache Spark and Apache Iceberg all shown as missing above an upgrade banner.
Captured 27 August 2026.

Teal's Match Score for this pairing opened at 9 percent, with Hard Skills at 3 out of 30 and Soft Skills at 0 out of 4. The chips shown as missing included Apache Kafka, Apache Spark and Apache Iceberg.

Hold that against the counts. Kafka appears twice in the resume. Spark appears five times. Iceberg appears three times. What the resume does not contain is the string Apache, which appears zero times, because engineers writing about their own work overwhelmingly write Kafka and Spark and Iceberg as bare tokens. The posting's keyword list carries the vendor prefix, the resume does not, and the full phrase match fails on all three.

This is not a mysterious scoring model. It is a literal string comparison doing exactly what a literal string comparison does, and it is worth being precise about the failure mode, because it is the same one that applicant tracking folklore has been warning about for a decade, only pointed in the opposite direction. The tool is not missing the candidate's skills because the candidate lacks them. It is missing them because of a one word prefix.

What the resume contains versus what the score saw: Kafka appears twice, Spark five times and Iceberg three times yet all three are marked missing, while Apache Flink, which never appears, becomes the only keyword marked present after one generated bullet. The match score read 9 percent before and 15 percent after.
Occurrence counts are from text extraction of the source file.

A user seeing 9 percent has been handed a real signal and a misleading one in the same number. The real signal is that the resume's wording does not mirror the posting's wording, which is fixable and worth fixing. The misleading signal is the implication that the underlying experience is absent. On this pairing it is abundantly present, ten occurrences of the three headline technologies across two pages, and the score cannot tell the difference.

Then the score paid out for a skill the resume never had

The score screen is only half the product. The other half is generation, and that is where the two halves come apart.

Teal's Write Bullet feature opens with a prompt that reads "Select up to 3 keywords to include in your bullet". This is a reasonable interface. It is also, structurally, a request for the user to choose what the resume should claim, from a list of what the job wants, with no reference back to what the resume already supports. We selected Apache Flink, the one term on the list that the control resume contains zero times.

It produced three bullets. All three asserted Flink experience. One read, in full: "Led the design and implementation of a high-performance data ingestion pipeline using Apache Flink, achieving a 50% reduction in processing latency within 6 months." Another shipped with a literal unfilled placeholder still in the text, "over X million events daily", which is a small thing but a telling one: the template is generating shape first and substance second, and nothing downstream caught that the shape had a hole in it.

Three generated bullets asserting Apache Flink experience, including one with the unfilled placeholder, over X million events daily.
Captured 27 August 2026.

Use Bullet placed the text into a plain edit box with a Save button. No warning, no confirmation step, no marker recording that this line came from a generator rather than from the candidate's history. Once saved, the line is indistinguishable from every other line on the document.

The same score screen after saving one generated bullet: 15 percent, Hard Skills 5 of 30, and Apache Flink now the only keyword shown as present.
Captured 27 August 2026.

Then the score moved. Match Score went from 9 percent to 15 percent. Hard Skills went from 3 out of 30 to 5 out of 30. Apache Flink became the only keyword on the screen rendered as present. Kafka, Spark and Iceberg, the three the candidate has actually used, were still marked missing.

That is the whole finding in one sentence: on this run, the only way to move the number was to add the one thing that was not true, and the tool's own display then certified it as the candidate's strongest match. The mechanism is not exotic. The matcher checks strings against the posting, the generator writes from the posting, and nothing in between checks the generator's output against the resume it is editing. Insertion and evidence produce the same reward because the scorer cannot distinguish them.

Where Teal is right, in its own words

Teal's documentation does not endorse any of this. It argues against it, clearly, in language this post cannot improve on.

A 75% match with real content beats a 95% match with fluff.

That line is from Teal's help center article on tailoring a resume and reading the Match Score [2], and it is the thesis of this entire benchmark stated by the company being benchmarked. The same article is direct about the trap we walked into on purpose: "If a job asks for 'Salesforce' and you've never used Salesforce, don't pretend you have." [2] A separate article on generated bullets warns that "our AI may generate metrics that aren't accurate to you" [3], which is a fair description of the 50 percent latency reduction and the placeholder events figure both.

So the gap here is not between what a company believes and what it says. It is between the help center and what the interface rewards at the point of use. The article explains the color coding as "Green = you have it. Yellow = toggle it on. Red = add it." [2] Add it is doing enormous work in that sentence. It means add the true thing you forgot to mention, and the product cannot tell that apart from add the thing you have never done, because at the moment the user clicks Save there is no check, no prompt, and no record of where the line came from. Documentation sits several screens away from the button. The score sits directly under it.

One more thing Teal did on the same day deserves saying, because it is the opposite behavior and it is genuinely good. Asked through its conversational job search for remote senior data platform roles combining lakehouse and streaming work at a high salary bar, it returned zero results and said so plainly, broadened the search once, still found nothing, and explained why. No filler listings, no adjacent roles dressed up as matches. The interface also carries a footer reading "Teal can make mistakes." A product willing to return an empty set and label its own uncertainty is a product with the right instinct available to it. That instinct just has not reached the scoring screen.

Jobright sends the phone to a desktop

Jobright's mobile feed with the banner, Your Resume needs improvement! Visit PC site.
Captured 27 August 2026.

Jobright's mobile web product does not offer resume tailoring. A banner sits persistently in the feed reading "Your Resume needs improvement! Visit PC site."

That ends the tailoring test for this benchmark, and it ends it cleanly rather than inconclusively. We could not exercise Jobright's tailoring on a phone because on a phone it does not exist. This post therefore makes no claim whatsoever about the quality of Jobright's desktop tailoring output. It was not run, so it is not graded.

What is worth reporting is the shape of the scoring, observed the same day on a job card for the same class of resume. Jobright showed an overall match of 82 percent, decomposed into Experience 100 percent, Skill 65 percent and Industry 44 percent. Set aside whether those specific numbers are right. The decomposition itself is a better artifact than a single literal keyword percentage, because it tells a candidate which dimension is weak and therefore what to do next. A lone 9 percent tells a candidate to panic. A 65 percent on skill next to a 100 percent on experience tells them to go fix their skills section, which is an action.

There is also a real cost to the mobile gap that has nothing to do with quality. A large share of job hunting now happens on a phone, in the gaps of a working day. A banner that says visit a desktop is a feature that, for that user, on that day, is simply not there.

Simplify puts tailoring behind Simplify+

The Simplify+ screen listing perfectly tailored resumes in one click among paid features.
Captured 27 August 2026.

On mobile on 27 August 2026, Simplify's tailoring sits behind Simplify+, whose feature list includes "Perfectly tailored resumes in 1-click". On desktop on 10 August 2026, on the free tier, clicking the generate button inside the tailor modal produced an upsell modal: no output, no partial output, no free tailored resume to inspect.

No subscription was purchased for this benchmark, so Simplify's tailored output was not tested. Its quality is not assessed here, at all, in either direction. That is a gap in the benchmark and it is stated as one.

What was observable on the free tier is worth more space than the paywall, because it is the best thing any of the three tools did on the day. Simplify's parse confirmation step handles uncertainty the way the rest of this category should. When it could not resolve one of the control resume's employers against its own database, it stopped and asked: "We couldn't find an exact match for Meridian Analytics in our system. Please confirm or add by searching it below."

Read that next to the Flink bullet. Both are moments where a system does not know something. One resolves the unknown by asking the person who does know. The other resolves it by generating a plausible sentence and moving on. The difference is not effort or intelligence, it is where the product decided the user belongs in the loop, and Simplify put them in the right place. It costs a tap. It buys a document whose employer names are correct.

Simplify's help center is similarly measured about scores. It advises getting to "around 70% keyword coverage or higher before applying" and then immediately adds that "A high score does not guarantee interviews, and not every keyword matters equally." [4] That second sentence is the part most score displays leave out.

The same resume in JobShifu

JobShifu's final step on the same control resume: Fit 80 percent, Readiness 51 to 76, and a notice that one baseline line had no confirmed evidence behind it and was left off.
Captured 27 August 2026. JobShifu is the publisher's product.

JobShifu is the publisher's own product and this section is labeled as such. It also ran against a different posting, a Senior Software Engineer, Evaluation Engine role at Vanta, using the same control resume. That is a real asymmetry in the benchmark and it is stated up front rather than buried: this section is not a like for like score comparison with the ServiceNow numbers above, and it should not be read as one.

On that run, Fit came out at 80 percent. Readiness moved from 51 to 76. Keyword coverage moved from 6 out of 15 to 10 out of 15, entirely from facts already confirmed in the user's Vault. And one line that had been in the baseline resume did not make it onto the tailored document, with a notice saying why: "1 baseline line(s) had no confirmed evidence behind them and were left off. Confirm the fact in your Vault to bring them back."

The stated gap was equally blunt: "No Cybersecurity and Compliance experience, often a listed qualification." The posting wants it, the Vault does not have it, and the product says so rather than writing a bullet about it.

The mechanism is one sentence, and the product states it on its first screen: "The Vault holds everything true about you. Tailoring picks what matches this job, and nothing lands on a resume that is not in here." That is the entire structural difference. The generator's input is the confirmed set, not the posting. Adding a keyword the user has never used is not a thing the interface can be talked into, because there is nothing to draw the line from. The cost, and it is a real cost, is the setup: the Vault has to be filled in before the tailoring is any good, and a notice about a dropped line is a chore the other products never hand you. More on how that works on /calibration and /resume-tailoring.

A five minute test you can run on any of them

None of this requires taking a benchmark's word for anything. The test that produced the finding in this post takes about five minutes and works on any tailoring tool, including JobShifu, on a free tier, without buying a thing.

Ask it to add a skill you do not have. Pick something adjacent and plausible, something a recruiter would believe of you and you have simply never done. Then watch what happens between clicking generate and the line landing on your document. Is there a prompt, a confirmation, a marker, anything at all? If the line saves silently, you have learned that the product will let the document say things you did not say, which is fine as long as you know it and are checking every line yourself.

Ask where a generated line came from. Click into a bullet the tool wrote and look for provenance: a source, a linked answer, a note about which of your inputs produced it. If the answer is that the line came from the job posting rather than from you, that tells you what the line is evidence of.

Watch which action moves the score. Add a true detail you had left out and note the change. Then add an untrue keyword and note that change. A scorer that rewards both equally is measuring string overlap with the posting, which is a useful thing to measure and a dangerous thing to optimize. Knowing which one you are looking at changes what you do with the number.

Method, data, and what this is not

The protocol, restated plainly:

  1. Build one control resume as a fixture: two pages, fictional identity, senior data platform engineer.
  2. Extract its text and count the terms that matter, before running anything.
  3. Run the same resume and one job posting through each tool on the same day, capturing every screen.
  4. Grade every tool verdict against the extracted counts, never against another screenshot.
  5. Publish the counts, the verdicts and the gaps between them.
Term Occurrences in the control resume
Kafka 2
Spark 5
Iceberg 3
Trino 1
EMR 2
Apache 0
Flink 0

Counts are by text extraction of the source file, re-verified 27 August 2026; the full record is in the data file listed under Sources [1].

What this is not: a census. It is one resume, one primary posting, four products, on one day, on a phone. It cannot tell you how any of these tools behave on a marketing resume, on a career changer, on a posting with vague requirements, or on desktop where two of them clearly put more of their product. Scoring models change without notice, and a run in November may produce different numbers than a run in August. Treat it as a probe with a published method: a thing designed to be repeated and disagreed with, not a leaderboard.

The dataset is published under CC BY 4.0 as tailor-benchmark-2026-08-27.json [1]. Cite it as: JobShifu, Resume tailoring tools benchmark, August 2026, with a link to this post.

If you re-run it and get something different, that is the useful outcome, and corrections are welcome. The point of publishing the fixture counts is that this post can be proven wrong with a text extractor.

Common questions

One resume and one job posting. Is that really enough to judge a tool?

No, and it is not offered as a judgment of the tools overall. It is a reproducible probe: a single controlled pairing, with the ground truth published as counts, run to see whether a specific mechanism holds. What it can support is a narrow claim about what happened on this pairing on this day. What it cannot support is a ranking, a quality verdict, or a statement about how any of these products behave across a population of resumes. The protocol and the data file are published precisely so that anyone can extend it with more resumes, more postings and more days, which is the only way the narrow claim becomes a general one.

Which of these should I actually use?

Genuinely depends on what you need, and three of the four have a clear best reader. Teal, for someone who wants a free job tracker that is good at being a tracker, plus a help center that is unusually candid about the limits of its own scoring. Jobright, for someone who wants a match score broken into parts they can act on, and who works on a desktop, and who values the metadata on its feed. Simplify, for someone who applies at volume and wants autofill, and who will appreciate that its parser asks for confirmation instead of guessing at an employer it cannot resolve. JobShifu, for someone who wants every line on the document traceable to something they confirmed, and who is willing to spend the setup time that requires.

If a low match score can be wrong, should I ignore scores entirely?

Not ignore, but read them for what they measure. A literal keyword score is a wording check, and a wording check is worth having: this benchmark's 9 percent correctly detected that a resume saying Kafka does not literally say Apache Kafka, which is a fixable difference and might matter to a filter somewhere. What it cannot do is tell you whether the underlying experience exists. Simplify's help center puts the caveat well, noting that a high score does not guarantee interviews and that not every keyword matters equally [4]. Teal's puts the trade the other way round, arguing that a 75 percent match with real content beats a 95 percent one without [2]. Both are right. Use the score to find missing wording for things you have done, and stop there.

Did you pay for any of the paid tiers?

No. Everything reported here happened on free tiers. That is a real limit on the Simplify result in particular: its tailoring sits behind Simplify+, the free tier returned an upsell instead of output, and so its tailored output is not assessed in this post in any way. Teal's score screen carries an upgrade banner as well, so it is possible that paid tiers behave differently from what is described here. Anyone with a subscription who re-runs the protocol will produce a strictly more complete result than this one, and the data file is structured to be extended with it.

Sources

  1. The benchmark data file (JobShifu's own measurement): tailor-benchmark-2026-08-27.json. Published 27 August 2026, CC BY 4.0.
  2. Teal help center, on tailoring a resume and the Match Score: help.tealhq.com/en/articles/14435726. Updated March 2026, read 4 August 2026.
  3. Teal help center, on AI generated resume bullets: help.tealhq.com/en/articles/9519206. Read 4 August 2026.
  4. Simplify help center, on keyword coverage and the resume score: help.simplify.jobs/articles/2175778. Read 4 August 2026.
← All posts