Behaviorally Anchored Rating Scales (BARS): A Practical Guide
Most performance rating scales don't actually rate performance. They rate the manager's mood on review day. "Meets expectations" tells an employee nothing they can act on, tells HR nothing defensible in an audit, and tells the business nothing about whether the person should be promoted, coached, or moved. Behaviorally Anchored Rating Scales (BARS) exist to fix that specific problem — by tying each point on a rating scale to a described behavior an evaluator can actually observe.
BARS is not new. Smith and Kendall published the method in the Journal of Applied Psychology in 1963. What has changed is the context: distributed teams, AI-assisted work, continuous feedback cycles, and workforce audits that require rubric-based evidence. Those shifts — distributed work, AI-assisted output, audit pressure — are exactly the conditions BARS was designed for, which is also why it is now more likely to be misapplied. This guide covers what BARS is, when it works, when it doesn't, and how to run it without spending 18 months building anchors nobody uses.
What is BARS?
BARS is a performance evaluation method that assigns numerical ratings to specific, observable behaviors rather than to abstract labels. Instead of asking a manager to decide whether an engineer is a "3" or a "4" on collaboration, BARS gives the manager two paragraphs describing what a 3 looks like and what a 4 looks like — and asks which one matches what the manager has actually seen.
A behavioral anchor is a short description of a specific, observable behavior that corresponds to one point on the rating scale — the evidence a rater matches against, rather than an abstract label.
The scale most commonly runs from 1 to 5 or 1 to 7, though 1 to 9 scales also appear in the literature. Each point has a behavioral anchor: a sentence or short paragraph describing the behavior that earns that rating. Anchors are best written by senior practitioners in the role, with HR facilitating rather than drafting, and they use the vocabulary of the work.
.png)
Example: collaboration on a distributed engineering team
| Rating | Behavioral anchor |
|---|---|
| 5 | Reviews teammates' pull requests within one business day, leaves comments that explain the "why" not just the "what," and pairs with junior engineers on request without being asked twice. |
| 4 | Reviews PRs within two business days, gives useful feedback, participates in design discussions with informed opinions. |
| 3 | Reviews PRs when tagged, shows up to design reviews, generally works well with the team. Rarely initiates cross-team contact. |
| 2 | PR reviews take three-plus days. Contributes in meetings only when directly asked. Has been in conflict with at least one teammate in the review period that required manager mediation. |
| 1 | Blocks other engineers by not reviewing PRs. Actively resists design changes proposed by peers. Team has raised concerns to the manager. |
Notice what the anchors do. They name specific observable actions — PR review time, meeting behavior, manager escalations. A manager filling this out has to think about the person, not about the label.
Why BARS matters more in 2026
Three things have pushed BARS back into relevance.
Defensibility. Promotion, RIF, and PIP decisions increasingly face legal and regulatory scrutiny. "The manager felt she wasn't a fit" doesn't hold up. A rubric-based rating with documented behavioral evidence does. This matters most in BFSI, healthcare, and any organization operating under adverse-impact analysis, but it applies broadly.
AI and calibration. As AI-assisted work reshapes what "good performance" looks like, generic ratings hide the shift. A developer rated "4 on productivity" in 2024 and "4 on productivity" in 2026 may be doing fundamentally different work. Behavioral anchors force the conversation about what actually changed.
Skills-based organizations. Companies moving toward skills-based hiring and internal mobility need consistent evaluation across managers, teams, and geographies. BARS is one of the few frameworks that survives that kind of scale without collapsing into "everyone's a 3."
What BARS actually delivers
- Reduced subjectivity in individual ratings. Two managers rating the same employee against the same anchors will land closer together than two managers using "meets/exceeds."
- Feedback employees can act on. "You're a 3 on collaboration" is useless. "You're a 3 because you review PRs when tagged but don't initiate design discussions — the difference to a 4 is showing up to design reviews with a point of view" is a coaching conversation.
- A calibration artifact. During level or promotion calibration, BARS anchors give the room something concrete to argue about. Arguments about anchors are productive; arguments about vibes are not.
- Documentation that survives audit. Behavioral evidence tied to a rubric is defensible in ways that narrative reviews are not.
Where BARS breaks
BARS isn't a general-purpose performance system, and pretending otherwise leads to the failed rollouts that gave the method its bad reputation in the 2000s.
Building anchors from scratch is slow. A proper job analysis for one role takes weeks. Doing it for 40 roles takes a year. Most organizations don't finish. The half-built BARS system in your HRIS from a 2019 initiative is a common artifact.
Anchors go stale. Behavior that defined "exceptional" in 2022 may be table stakes in 2026. If nobody owns updating the anchors annually, the scale drifts out of usefulness within two review cycles.
BARS misses the shape of some roles. For research scientists, staff-plus engineers, or highly ambiguous roles where the work itself is defining the standard, behavioral anchors written last year won't capture this year's work. BARS handles predictable roles well and novel roles poorly.
Manager training is non-negotiable, and often skipped. A BARS system where managers haven't been trained on how to apply anchors — and calibrated against each other — degrades to the same subjective ratings BARS was meant to fix, just with more paperwork.
How to implement BARS without the 18-month project
The failed BARS rollouts almost always share one pattern: too much scope, too early. Here is a version that, in practice, tends to work.
Step 1: Pick three to five competencies, not fifteen
Start with the competencies that actually differentiate performance in the role. For a software engineer, that's usually some combination of technical delivery, code quality, collaboration, and ownership. For a customer support lead, it's response quality, escalation judgment, and coaching. Resist the urge to cover every dimension of the job.
Step 2: Let the people who do the work write the anchors
Anchors written without input from the role tend to sound generic. Anchors written by senior practitioners read like the job. Run a two-hour workshop with three to five senior people in the role, give them a rating scale, and ask them to describe what each rating looks like in specific observable behavior. Edit for clarity, but do not rewrite for polish.
Step 3: Pilot with one team before rolling out
One team, one review cycle, one competency set. Collect feedback from both managers and employees. The most useful signal: did the anchors help the manager write feedback the employee could act on? If not, the anchors are too abstract.
Step 4: Calibrate before you scale
Before rolling out to more teams, run a calibration session where managers rate the same three anonymized employees using the same anchors. The gap between the highest and lowest rating tells you whether the anchors are specific enough. As a working heuristic, a gap of two or more points suggests the anchors aren't specific enough yet.
Step 5: Own the update cycle
Anchors need annual review. Assign one person accountability. If nobody owns it, this fails.
What BARS doesn't do — and what to pair it with
BARS rates observable behavior in defined competencies. It doesn't measure business outcomes, doesn't capture growth trajectory, and doesn't tell you whether someone is ready to lead a team they haven't led yet. Most organizations that use BARS well pair it with:
- OKRs or goal frameworks for outcome measurement
- 360-degree feedback for cross-team collaboration signal
- How skills-based hiring works for validated capability, especially in technical roles
On the last point: if you're evaluating engineers on technical delivery competencies, BARS captures the behaviors (code review quality, design contribution) but not the underlying skill level. Pairing behavioral evaluation with structured skills validation gives a fuller picture. HackerEarth's skill assessments measure the underlying capability directly, so a BARS rating like "she's a 4 on technical delivery" can be checked against an independent skill signal — useful for calibration and internal mobility decisions.
.png)
FAQ
Is BARS worth the effort for a company under 500 employees?
Usually not for the full framework. Under 500 employees, the calibration problem BARS solves is manageable through informal manager conversations. Start with structured competency definitions and a simple 1–5 scale; graduate to full behavioral anchors when you have enough managers that calibration drift becomes a real problem.
Can BARS be applied to AI-assisted roles where the work itself is changing fast?
Partially. BARS works for the human-judgment layer of AI-assisted work — how someone evaluates AI output, escalates edge cases, or coaches teammates through unfamiliar workflows. It works less well for the technical execution layer, where "good" is moving too fast for annual anchor updates. For those roles, pair BARS with more frequent skill re-validation.
How does BARS handle bias?
BARS reduces rater inconsistency but leaves demographic bias largely intact — managers still know who they are rating. It is not a substitute for blind evaluation where blind evaluation is possible. Calibration sessions, where multiple managers rate the same anonymized employees against the same anchors, are the main mitigation for the bias BARS does not solve on its own.
How do you write a good behavioral anchor?
The hardest test is the swap test: if you can swap adjacent anchors on the scale without changing what a rater would decide, the anchors are not discriminating between levels. Good anchors also fail loudly — a rater should be able to say "I have not seen this behavior" rather than defaulting to the middle. If most of your ratings cluster at 3, the anchors at 2 and 4 are probably too demanding or too vague, not the raters too generous.
Next steps
BARS earns its keep when performance conversations shift from "I feel like you're a 3" to "here's what a 3 looks like, here's what a 4 looks like, and here's where you actually are." That shift is worth doing — but only if the anchors are built by practitioners, updated annually, and paired with the manager training that makes them stick.
If your organization is building toward skills-based evaluation, book a demo of HackerEarth Assessments to see how skills validation pairs with behavioral evaluation in your performance framework.



