“Like A Coach Explaining the Why, Not a Rigid Grader”

That’s a real line from a career center’s evaluation criteria, sent to our team ahead of a demo. That line doesn’t describe a feature request — it’s the criteria for what good guidance is supposed to feel like. That criteria came from a team running a deliberate, multi-vendor review. You may be familiar with how those look: several platforms on the list, requirements written out in advance, with a shortlist to follow.

Not every center gets there the same way, though. Ask five career centers how they picked their last platform and you’ll get five different stories. One already knew who they wanted and was collecting two more quotes because procurement required it, not because they were undecided. Another let the whole advising team weigh in, because the director felt every team member needed to buy-in on the tool they would be using. A third is mid-trial right now, running a scorecard nobody called a “framework” but that quietly asks the same questions anyway.

Different processes, different timelines, different politics. Some evaluations are objective and rigorous, some are just going through the motions to get where they want to go. But underneath all of them, the same three questions end up deciding it, whether anyone wrote them down or not. The center that asked for “a coach explaining the why, not a rigid grader” had simply already put one of those questions into words.

The reality is that most evaluations never get that specific. They begin and end with a feature list instead — a spec sheet where every capability gets an equal checkmark. That’s a reasonable place to start, but it’s also where the real answer is hiding. A checkmark can’t tell you what happens next:

Looked fine on paperWhat actually happened
AI-powered resume match-rate scoringThe score measured against a school’s benchmark rubric the vendor programmed into the tool — not against what employers were actually rewarding — so a high score didn’t reliably predict whether a resume would clear a real screen.
Full platform license, unlimited user seats
One center found users hadn’t used the tool at all over the past year. Another had fewer than 200 active users despite an institution-wide contract.
Strong registration and attendance at a tool rollout session
Students showed up once. The tool never became part of their routine afterward — usage dropped off within a semester because it didn’t fit how they actually worked.
A usage dashboard included in the platformDidn’t stop “how many students actually used it this semester?” from being the single most common question career centers ask at their own renewal.

Every one of those checkmarks was true. None of them was proof the tool actually did its job. That’s the gap a feature list can’t see — and it’s exactly what the three filters below are built to catch.

The alternative isn’t a longer checklist. It’s a smaller number of harder questions, run in a specific order, that a feature list can’t answer for you. Instead of grading tools, run every vendor through three filters. A vendor that doesn’t clear all three isn’t a fit — no matter how strong the demo was.

Filter 1: The Standard Filter

Does this measure guidance against what employers actually reward — or against an internal sense of “good enough”?

A lot of career tools grade against a static rubric: formatting rules, generic best practices, a checklist someone wrote once and never revisited. That’s not the same as measuring against what’s actually getting candidates hired right now — a standard that shifts as hiring practices and applicant tracking systems change underneath it.

The tell is whether a tool can explain itself. If it can’t tell you why its guidance is right — not just assert that it is — it isn’t applying a standard. It’s applying an opinion, and opinions don’t hold up when an applicant gets filtered out by a system nobody in the room designed the guidance around.

By applying the filter here, you can reveal how some tools use a scoring approach that perpetuate the issue, such as “benchmarking”. You upload a set of resumes considered good, and the tool learns to score new resumes against that set. It sounds reasonable until you notice what it actually measures: if the benchmark is based on uploads of only business-major resumes, every applicant who isn’t a business major gets scored against a standard that was never built for them. The tool isn’t wrong, per se — it’s just not answering the question that’s being asked. It can tell you “this looks like the resumes we uploaded.” It can’t tell you “this is what gets someone hired.”

That’s what the criteria from the career center we introduced in the opening was looking for. Their evaluation criteria had a second ask sitting right next to the “coach, not grader” line: “actionable learning over automation”. They were looking for a platform that teaches a user how to craft a stronger bullet point rather than one that quietly rewrites it for them. Underneath both asks is the same test: can the tool show its work, or does it just hand over an answer? A platform that only produces the fix, without explaining why the fix is right, fails the Standard Filter even when the fix itself is correct — because the next user who passes their resume through it gets no smarter for it, and the resume doesn’t get any better.

This filter matters most at the start of an evaluation, because it’s the easiest one to skip past on trust. A polished interface and a confident sales pitch can stand in for rigor if you don’t ask directly: measured against what, exactly, and how do you know that’s still true?

Filter 2: The Consistency Filter

Does the benefit reach every student — including the ones who never book an appointment — or only the ones who show up?

This is the filter most evaluations skip entirely, because it’s not really about the tool in isolation. It’s about your team’s capacity, and whether the tool deepens or widens that capacity or just adds a nicer diving board for the students who were already hopping into the pool.

One of the reasons this filter is so important is because it’s multi-layered. It’s about both the advisors and the students.

With the wrong tool, the effectiveness of the tool can still depend on which advisor a user happens to sit down with. ATS system and employer hiring standards shift constantly, while every career field has evolving standards. And sometimes the most knowledgable person on staff moves on or goes on leave, and that institutional knowledge leaves the building with them. Most career centers don’t have a shared, current source of truth for what “good” looks like right now. Therefore, guidance quality ends up depending on how current any individual advisor’s own knowledge happens to be. A newer advisor gives different guidance than a ten-year veteran, not because either one is wrong, but because there’s no shared standard underneath either of them.

A tool that actually clears the Consistency Filter closes that gap: every advisor on staff, regardless of their own ATS fluency or how current their hiring-trend knowledge is, ends up giving guidance built on the same standard. That’s a systems gap, not a competence gap — the difference between “our guidance is only as good as whichever advisor happens to be in the room” and “our guidance is consistent no matter who’s in the room.”

The second layer is a question, for those in higher ed specifically: if a student never sets foot in your office, do they still get the benefit of your team’s standard? If the answer is no, you haven’t solved the consistency problem — you’ve handed your most-engaged students a better experience and left everyone else exactly where they started. It’s the same reason a fully licensed platform can sit at under 200 active users despite full institutional buy-in: the seats existed, the reach didn’t.

Sometimes the mechanism is more specific than simple non-use, too. One school’s own internal assessment named the exact moment reach breaks down: students disengaging the instant they hit a rigid scoring system, which is why the team went looking for a gentler on-ramp for students starting from zero. That’s the same failure the rollout-session example pointed at earlier — attendance and access aren’t the same as fit. A consistency failure doesn’t always look like “nobody logged in.” Sometimes it looks like “they logged in once, hit something that didn’t work for them, and never came back” — a harder failure to catch, and an easier one to miss entirely.

Filter 3: The Proof Filter

Can you produce a number for your VP or Provost that ties this to an outcome — not just an activity count?

Login counts and usage stats are activity. They tell you a tool got used; they don’t tell you it worked. The Proof Filter asks for something harder: a defensible line between “we invested in this” and “this changed a placement number, an interview rate, or an employment outcome.”

This isn’t only a pre-purchase filter. It’s often the difference between a renewal and a churn conversation. The centers most likely to walk away from a tool aren’t usually the ones with a bad product experience — they’re the ones who can’t answer “did this work?” when their own leadership asks, because “how many students actually used it this semester” is already the most common question career centers face at their own renewal. If proof isn’t built in from day one, it isn’t a gap you close later. It’s a gap that grows every semester you don’t have the answer.

The proof bar is rarely abstract, either. One institution’s ask ahead of a trial was specific down to the field: could the platform export data aligned with the student identifiers their office already used for cross-campus reporting, so it could flow straight into the dashboard leadership already looked at. Proof that doesn’t fit into the reporting your institution already runs is proof nobody but you will ever see.

The Filters Aren’t Invented — They’re Already How Buyers Think

One school’s own trial scorecard split almost exactly this way without ever naming it that: their internal criteria bucketed into feedback quality, coaching impact and user engagement, and data and reporting — Standard, Consistency, and Proof, just in their own words instead of a framework’s. The three filters aren’t a new invention. They’re a name for questions serious buyers already ask themselves.

One Tool, One Job

A tool can legitimately clear all three filters for a narrow job — interview practice, say — and still not be the whole answer for everything your center needs. That’s fine. The failure mode isn’t “the tool does one thing well.” It’s assuming one tool has to do everything, which is a different mistake with a different fix.

The Consolidation Trap

Career centers are actively trying to reduce the number of platforms students are asked to use, and that pressure can override all three filters at once. A tool that passes Standard, Consistency, and Proof individually can still lose if it’s evaluated as “one more login” instead of a layer on what’s already in place. One evaluation note put the pressure plainly: the team was testing whether a platform could “credibly replace” two existing tools “so we can simplify our stack and justify spend.” The deciding question there was never just “is this good” — it was “does this let us subtract something,” which is a different bar entirely. If you’re running these filters against a shortlist, it’s worth asking a fourth, adjacent question: does this replace something already in the stack, or does it sit awkwardly next to everything else students are already being asked to juggle?

Why the Order Matters

Most vendor pitches lead with Filter 3 — their ROI stats, their case studies, their logo wall. That’s backwards. A proof point from a tool that never solved the standard or consistency problem underneath it isn’t proof of value. It’s proof that some users had a good experience, which is a much smaller claim than it’s usually dressed up to be.

Run the filters in order. A vendor that skips straight to proof is usually hoping you won’t ask about the first two — and if you don’t, you’re the one who has to explain the gap later, not them.

The Framework Holds Either Way

None of this is an argument for a specific tool. It’s an argument for a specific order of questions, and it holds regardless of what ends up on your shortlist.

It’s also the lens we try to hold ourselves to. If you want to see how we think about applying it in practice — including the market data and vendor scoring we used to stress-test our own thinking — the Career Tech Buyer’s Guide walks through it in more depth.

If you’re heading into a budget conversation this cycle, running your shortlist through these three filters first might save you a few false starts.

Click to rate this article
[Total: 0 Average: 0]