CASE STUDY — ELSEVIER

Libra — Designing Trust into AI Evaluation at Scale

How UX design turned an internal engineering platform into Elsevier's system of record for expert AI evaluation

From raw infrastructure to a guided, role-based evaluation experience — 100+ design-driven tickets shipped across 15 months, protecting the integrity of the data that validates Elsevier's AI products.

RoleUX Designer, end to end
Duration15 months
PersonasManagers, Evaluators, Admin
Shipped80+ of ~109 tickets
01

Overview

Libra (SME Evaluation Tool) is Elsevier's internal platform for structured expert evaluation of AI-generated content. Subject Matter Experts — clinicians, researchers, domain specialists — score AI responses from products like ClinicalKey AI, ScienceDirect AI, and Sherpath AI against configurable quality metrics. Their judgments feed directly into product decisions about whether AI features are safe and good enough to ship.

I joined as the UX Designer responsible for the product's end-to-end experience, working across three personas (Evaluation Managers, Evaluators/SMEs, and Panel/Admin roles) and every major surface: onboarding, evaluation setup, the evaluator console, the manager console, datasets, teams, and metrics. Over 15 months, the design practice on Libra evolved from producing screens for individual features into a systematic discipline — role-based interface audits, in-product guidance strategy, and design-led protection of data integrity.

02

Business Context

Elsevier's AI products operate in high-stakes domains: clinical decision support, scientific research, medical education. Before these products reach users, their outputs must be validated by human experts. Libra is where that validation happens — it is the productionized version of an earlier proof of concept, built to support SME-based evaluation workflows, intelligent query routing, and evaluation analytics.

Three business realities shaped the design work:

Expert time is scarce and expensive. SMEs are practicing physicians and researchers. Every minute lost to confusing UI is a minute of clinical expertise wasted — and evaluator fatigue directly degrades rating quality.

The output is data, and the data must be trustworthy. An evaluation platform that allows accidental submissions, ambiguous states, or misinterpreted deadlines doesn't just frustrate users — it contaminates the evidence base used to approve AI features.

Adoption depended on self-sufficiency. The SME Operations team was spending significant effort on live training sessions, documentation maintenance, and answering repetitive "how do I do X?" questions. As the evaluator community grew, this model could not scale. The product had to teach itself.

03

Why This Was Hard

A product born from engineering, not experience.

Libra's first months were infrastructure topology, database design, and API contracts. UX arrived to a system where technical capability existed but usability was an afterthought — and had to retrofit coherence without blocking delivery.

Three personas, one interface, overlapping permissions.

An Admin can act as Manager, Panel, Evaluator, or staff; a Manager can act as Panel or Evaluator. Designing a single interface that communicates "your highest role" without implying profile-switching — while showing each person only what their current task requires — was a recurring, genuinely difficult information-architecture problem.

Extreme content density.

Evaluator sessions involved AI summaries up to 15 pages long, rated against up to 17–18 metrics, plus reference-level and field-specific scoring. The core design tension: how do you keep an expert oriented, efficient, and accurate inside that much content?

Conceptual coupling in the domain model.

Datasets and evaluations were tightly entangled in the original flow, confusing ownership, lifecycle, and reuse. Untangling them was as much a product-modeling exercise as a UI one.

A moving target.

Feedback arrived continuously — UAT sessions, module-wise workshops, stakeholder emails — while the platform was actively migrating real evaluation programs onto it. Design had to triage, systematize, and ship simultaneously.

04

My Role and Ownership

I owned UX design for Libra end to end: interaction design, UX auditing, UX writing, and design communication with Product Managers and the broader team. In practice this meant:

Sole designer on the majority of design tickets from late 2025 onward, covering the evaluation setup wizard, dataset flows, evaluator console, manager console, teams, and role systems.

Driving the audit program — I organized the legacy design library against the live Libra interface (including PRD mockups), then led role-based audits of the Manager profile (three parts) and Evaluator profile (four parts), converting scattered feedback into a structured, prioritized backlog.

Proposing strategy, not just executing tickets. The Coach Marks & Guided Onboarding initiative began as my written proposal, arguing from operational cost data that in-product guidance would scale where live training could not. It was prioritized as High and shipped.

Running a UX audit of the entire Jira backlog, surfacing UX-relevant work across 11 product areas and organizing it into an Assumption Matrix (risk × knowledge quadrants) to give PMs a shared prioritization language.

Owning critical-path design. Eight of the tickets I delivered were classified Critical — including empty states, closed-evaluation integrity, conditional submission, and reviewer identification — meaning design sat directly on the product's risk register, not its polish list.

05

Approach and Key Decisions

Decision 1

Separate the dataset from the evaluation.

I championed treating datasets as reusable assets and evaluations as configuration layers applied to them: create a dataset once (with rich, standardized metadata — product, version, model, stage, generation date), then select it when creating evaluations. This reduced setup errors, enabled reuse across evaluation programs, and simplified the wizard's cognitive load.

Decision 2

Design by role, not by screen.

Instead of patching individual pages, I audited each persona's complete journey and defined role-based interfaces (RBAC-aligned). The role-switching design and multi-role display work resolved a Critical confusion — users believed they were supposed to click to switch profiles — by clearly communicating capability rather than implying action.

Decision 3

Make the product teach itself.

I designed a layered guidance system: a welcome tour modal, a 9-step coach-mark tour through the evaluator console (start action → evaluation queue → query panel → response panel → core metrics → panel collapse/expand → field-specific metrics → submit), contextual coach marks for feature announcements, inline tooltips across the setup wizard, and calibrated empty states. Copy tone was contextual by design — "All caught up" when a manager's review queue is genuinely clear, neutral "No results found" for filter combinations — because an empty state can be a reward or a dead end, and the interface should know the difference.

Decision 4

Treat integrity failures as UX-critical.

When a CLOSED evaluation still allowed full interaction, or "Look Inside" silently auto-assigned items and rendered an editable submission UI, I treated these as first-order design problems: state must be legible, and destructive or binding actions must be deliberate. Similarly, surfacing the implicit UTC timezone on evaluation start/end dates eliminated a failure mode where evaluations appeared to open late in Asia and close early in the US.

Decision 5

Reduce evaluator fatigue structurally.

For the dense evaluation screen, I redesigned the layout around a three-column model with collapsible panels, disambiguated iconography (panel controls vs. dropdowns vs. reference navigation), explicit labeling of the metrics area, a guided response-metrics-then-reference-metrics flow, and progress visibility — directly addressing workshop findings about 15-page summaries and 18-metric scoring.

Decision 6

Design the manager's trust loop.

The manager console gained legible query-status legends, filter-respecting exports, review-status filtering, and inline reviewer-role identification chips — so a manager can always answer "what needs my attention, who reviewed what, and can I defend these results?"

06

Design System Impact

Libra's visual language — clean white surfaces, the blue accent family (#2557C5, #185FA5), pill badges — was consolidated into a consistent, reusable vocabulary as the work matured:

Componentized status communication: a unified legend/badge system for query states (completed, in progress, disagreement, cannot rate, qualitative feedback) with defined visual hierarchy for multi-status combinations, built to scale as new states are added.

Standardized role identity: inline role chips with documented edge cases, reused across the results table, team cards, and profile surfaces.

Pattern library for guidance: coach mark anatomy, tour modals, and feature-announcement banners defined once and applied across surfaces — the Comparative Evaluation work explicitly shipped on existing design-system tokens.

Form and wizard conventions: grouped toggles with structured warnings, card-radio selectors with disabled-plus-tooltip states, and consistent stepper behavior across the setup wizard.

State coverage as a standard: empty, loading, error, closed, and unassigned states became mandatory deliverables in every design ticket, not afterthoughts.

The audit artifacts themselves — the old-designs-vs-Libra mapping and the Assumption Matrix — became living team infrastructure, giving PMs and engineers a traceable link between design intent and shipped product.

07

Outcomes

~109 UX/design-related tickets across 15 months, 80+ shipped, spanning every product area — with 8 Critical-priority design tickets delivered, placing design on the product's critical path rather than its cosmetic layer.

Data-integrity failure modes eliminated by design: closed evaluations now communicate their state and block accidental edits; silent auto-assignment was removed; timezone ambiguity on deadlines was resolved. Each of these protected the validity of evaluation data feeding AI product decisions.

A scalable onboarding model. The coach-mark system replaced dependence on live training with in-context learning — designed explicitly to reduce onboarding time, lower the support burden on SME Operations, and hold up as the evaluator community grows.

Faster, safer evaluation setup: the dataset/evaluation separation, standardized metadata, wizard restructuring, and workflow-language rewrites (replacing terms stakeholders flagged as incomprehensible in UAT) reduced configuration errors and repeated support questions.

Manager confidence in results: review-status filtering, reviewer-role identification, and filter-aware exports gave managers end-to-end traceability over evaluation quality control.

A design practice that compounds. New workstreams — Endpoints, LLM-as-Judge, Comparative and Parallel evaluation, search-quality dashboards — now arrive to established patterns, audit processes, and a shared prioritization framework, lowering the design cost of every future feature.

08

Key Learnings

In evaluation platforms, UX is a data-quality discipline. The most consequential design work wasn't visual — it was making system state legible and binding actions deliberate. A confusing interface here doesn't just annoy users; it corrupts the evidence used to ship AI to clinicians.

Audit before you redesign. Systematically mapping legacy designs against the live product, then auditing persona by persona, converted a fog of feedback into a prioritized, defensible backlog — and earned design a seat in roadmap conversations.

Guidance scales; training doesn't. The strongest argument for the coach-mark investment wasn't aesthetic. It was operational: quantifying the recurring cost of human-led onboarding made in-product guidance a business case, not a design preference.

Copy is interface. Renaming workflows that stakeholders couldn't parse, and calibrating empty-state tone to context, produced some of the highest clarity-per-effort wins in the project.

Design the model, not just the screens. The dataset/evaluation separation showed that the deepest UX improvements sometimes live in the product's conceptual architecture — and that designers should argue for those changes explicitly.

09

Closing Statement

Libra taught me what design looks like when the user interface is also a scientific instrument. Every legend, empty state, and confirmation modal exists to protect one thing: the trustworthiness of expert judgment about AI. Over 15 months, I helped transform an engineering-first platform into a product that onboards its own users, respects its experts' time, and defends the integrity of its own data — and built the design system and audit practice that will keep it doing so as it scales.