How to Evaluate AI Before & After Simulations
A practical framework for judging identity preservation, treatment fidelity, unrelated-feature drift, failure handling, disclosure, consent, retention, and consultation usefulness.
The right question is not “Does this image look impressive?” It is “Does this system create a useful, bounded visual reference while preserving the person, changing only the intended treatment area, handling failures honestly, and protecting the photo throughout the consultation workflow?”
An AI before-and-after image can make an abstract aesthetic preference easier to discuss. That value also creates risk: a polished image can be mistaken for a clinical prediction, and a subtle identity change can make a result feel more appealing for the wrong reason. A serious evaluation therefore has to cover the image, the product workflow, and the provider’s use of it.
1. Define the job before scoring the image
“Before-and-after simulator” can refer to several different products. A consumer website tool helps a person explore a visual idea before speaking with a practice. A provider consultation tool helps a licensed professional clarify or present a preference during assessment. A surgical planning system may model anatomy in three dimensions. Those jobs should not share one undifferentiated accuracy claim.
Write the intended job in one sentence before the evaluation begins. For example:
- Public discovery: Create a conservative, clearly labeled illustration that helps a visitor decide whether to request a consultation.
- Provider consultation: Give the patient and provider a visual reference they can refine, contextualize, or reject after a real assessment.
- Not the job: Diagnose a concern, determine candidacy, select a product or dose, choose a technique, model healing, or guarantee an outcome.
This boundary changes the rubric. A public tool should be judged heavily on disclosure, permission, failure handling, and safe routing. A provider tool also needs role access, presentation controls, case context, and a clear separation between the simulation and the actual plan.
2. Replace one accuracy number with a scorecard
A single percentage hides the failure modes that matter. A model could earn a high average resemblance score while changing teeth in ten percent of portraits. It could generate photorealistic skin while editing outside the requested area. It could look plausible on one regeneration and drift on the next.
| Dimension | Question to score | Example failure |
|---|---|---|
| Identity preservation | Is this unmistakably the same person? | Face shape, age, ethnicity, eye shape, or distinctive features change. |
| Treatment fidelity | Is the visible change confined to the intended area and direction? | A lip preview changes the whole lower face or ignores the selected intensity. |
| Unrelated-feature stability | Did eyes, teeth, nose, hair, ears, jewelry, clothing, lighting, and background stay stable? | Whiter teeth or smoother skin make the “after” more appealing independently of treatment. |
| Texture and anatomy plausibility | Does the illustration retain believable texture, edges, symmetry, and proportions? | Plastic skin, doubled contours, warped jewelry, or erased folds. |
| Repeatability | Do repeated generations stay within the same visual direction? | The identity or apparent treatment magnitude changes widely on each run. |
| Failure honesty | Does the system reject an unusable source instead of inventing confidence? | A profile view or obstructed face produces a polished but unreliable edit. |
| Consultation usefulness | Does the image help a provider clarify the preference or correct expectations? | The image is attractive but too ambiguous or exaggerated to discuss responsibly. |
Use a small anchored scale
A four-point scale is easier to calibrate than a vague ten-point score:
- 0 — unusable: Material identity change, major artifact, wrong target, or unsafe presentation.
- 1 — weak: Noticeable drift or ambiguity that undermines consultation use.
- 2 — usable with explanation: The intended direction is clear; minor issues remain visible.
- 3 — strong: Identity and unrelated features remain stable; the bounded change is clear and plausible as an illustration.
Keep each dimension separate. Do not let a beautiful overall image cancel out a zero for identity preservation or disclosure.
3. Build a representative portrait set
A vendor-selected gallery is a demonstration, not an evaluation. Build a set that represents the people, photos, and treatment questions the product will actually encounter. Obtain the necessary permissions for any evaluation image and keep test data out of public materials.
A useful internal screen can start with 40–60 consented or fully synthetic portraits, then expand as the product reaches more environments. Include variation across:
- skin tone, age range, face shape, gender presentation, and visible skin texture;
- glasses, makeup, facial hair, jewelry, hairstyles, and partial occlusion;
- even light, side light, overhead light, shadows, and common phone-camera quality;
- front-facing usable sources plus deliberately unsuitable profiles, cropped faces, group photos, filters, and heavy blur;
- each supported treatment area and every customer-facing intensity level.
Keep a stable core set for regression testing. Add a separate challenge set when a real failure appears. This distinguishes “the current release improved” from “the evaluator happened to choose easier photos.” Record the model and prompt version, product configuration, date, and whether a generation succeeded, failed, timed out, or was rejected before generation.
4. Run a reproducible image test
- Freeze inputs. Use the same source image, treatment key, intensity, and product version for every comparison.
- Blind the review. Hide vendor and model labels when reviewers score image quality.
- Use at least two reviewers. Include a trained product reviewer and a licensed provider familiar with the treatment area. Resolve large disagreements with written reasoning.
- Generate repeats. Run more than one output for a subset of cases to measure variation rather than judging a lucky result.
- Record rejection quality. A correct “please retake this photo” is better than a confident artifact.
- Preserve failures. Do not remove unattractive outputs from the denominator. Separate provider errors, validation failures, timeouts, and user-correctable source problems.
- Publish the denominator. If results are shared, state the number of portraits, treatment mix, successful generations, rejected inputs, failures, review method, date, and limitations.
Inspect identity and drift systematically
Compare the source and simulation region by region. Start outside the requested treatment area: hairline, brows, eyes, nose, teeth, ears, jaw edge, jewelry, clothing, and background. Then inspect skin tone, lighting direction, camera perspective, and overall age impression. Only after those checks should the reviewer score the intended change.
This order matters. If reviewers first admire the treatment result, they are more likely to overlook unrelated beautification. A before-and-after tool should not quietly turn into a whole-face filter.
5. Audit the clinical and communication boundary
An image can be visually useful without being clinically predictive. A single portrait does not establish three-dimensional anatomy, muscle movement, medical history, contraindications, product behavior, technique, swelling, healing, or the person’s final result.
Check the complete customer journey for consistent language:
- before the user uploads a photo;
- next to the generated result;
- inside any comparison slider, saved image, email, or shared link;
- in the provider presentation or case view;
- at the consultation and booking handoff.
The disclosure should be specific: “illustrative simulation,” “not a clinical prediction,” “individual results vary,” and “licensed-provider assessment required.” “AI generated” alone does not explain the limitation.
6. Treat consent, privacy, retention, and deletion as product quality
Facial photos and treatment interests deserve a stricter review than ordinary marketing analytics. Ask the vendor to walk through the data lifecycle on screen and in writing.
Before processing
- What exact permission does the person see, and is it required before capture or upload?
- Are photo processing and consultation support clearly described?
- Is promotional contact separate, optional, and unchecked?
- Does the permission authorize model training? If not, is that exclusion explicit?
During storage and access
- Where are source and generated images stored, and which subprocessors receive them?
- Are images private by default? Are access links bounded and revocable?
- Can staff see only the correct location or tenant?
- Are raw image bytes, unrestricted URLs, and contact details excluded from analytics and logs?
At the end of the lifecycle
- Can a person request deletion without navigating an obscure support process?
- What is the maximum retention period, and what event starts or extends it?
- Does deletion cover source images, generated images, derived artifacts, and cached copies?
- Can the practice verify deletion and retention behavior before launch?
For Eva’s current public-flow boundary, read the Photo & AI Notice. The public flow requires a versioned photo-and-consultation permission, excludes model training, separates promotional consent, supports deletion, and defines a maximum retention boundary. The provider planner is a distinct workflow and should not be described as having identical controls until parity work is complete.
7. Observe a real consultation workflow
Image scores alone cannot show whether the product improves a conversation. Run a structured pilot in which providers use the image as one input—not as the plan. After each consultation, ask:
- Did the image help the person express a preference more clearly?
- Did it expose an expectation the provider needed to correct?
- Could the provider explain why the real plan might differ?
- Did the simulation save time, add time, or simply shift the conversation?
- Was the provider comfortable presenting, rejecting, or regenerating it?
- Did the system preserve the correct person, location, permission, and follow-up context?
Separate product events from business outcomes. Preview completion, consultation request, staff contact, booked appointment, attendance, treatment, and attributable revenue are different states. Report the time window and attribution rule. Do not turn a handful of anecdotes into a conversion-rate claim.
8. Use the same script in every vendor demo
Bring your evaluation set and require the vendor to run it live. Ask to see a good source, a poor source, a provider rejection, a regenerate path, a deletion request, and a multi-location access boundary. Then request the current written terms for subprocessors, retention, training, incident response, included usage, failure credits, onboarding, and renewal.
Compare the complete offer. A promotional monthly price is not enough. Record:
- the exact product or module included;
- face, body, 2D, 3D, and treatment scope;
- public website capture versus provider-only use;
- branding, embedding, routing, and booking handoff;
- successful generations included and how failures count;
- setup fees, contract term, promotional period, and renewal price;
- consent, retention, deletion, and training language.
Eva’s EntityMed comparison shows this dated, scope-first approach. Unknown facts are marked for buyer verification rather than filled with guesses.
Red flags that should stop a launch
- A universal accuracy claim without a defined test set, rubric, denominator, and date.
- Results presented as likely, expected, predicted, or guaranteed clinical outcomes.
- Material changes outside the requested treatment area.
- No graceful rejection for unusable photos.
- Photo processing before meaningful permission.
- Training rights bundled into required processing permission.
- No concrete maximum retention period or deletion path.
- Public or guessable image URLs, raw photos in logs, or cross-location access.
- Provider presentations that drop the simulation label.
- Pricing comparisons that omit promotion length, usage, onboarding, or renewal.
A compact decision template
Before approving a pilot, require one page that records the intended job, supported treatment scope, evaluation-set composition, rubric and pass thresholds, known failure modes, disclosure text, permission version, subprocessors, maximum retention, deletion test, staff roles, included usage, contract term, and the person accountable for launch review.
Then keep the document alive. Re-run the stable evaluation set whenever the model, prompt, treatment definition, image validation, consent text, storage path, or presentation surface changes. An AI before-and-after system is not “accurate” once and forever; it is a versioned product whose quality and safety need continuing evidence.
Apply the framework to Eva
You can inspect the consumer job directly in Eva’s free public preview, explore treatment-specific boundaries in the Botox simulator and lip filler simulator, or review the practice workflow. Generated examples are labeled as synthetic demonstrations or illustrative simulations, never patient outcomes.
If you are evaluating Eva for a practice, ask the team to walk through the same checklist in your demo. A useful pilot should make the boundary easier to inspect—not ask you to trust a gallery.
Frequently asked.
- There is no meaningful universal accuracy percentage. A useful evaluation separates identity preservation, treatment-area fidelity, unrelated-feature drift, realism, repeatability, failure handling, and provider-rated consultation usefulness on a defined portrait set.
The rest of this series.
Eva AI Team
Product and editorial team
The team documents Eva's AI before-and-after product, image-quality boundaries, consent and privacy controls, provider review, and aesthetic consultation workflows.
Published
Continue reading.
- IAI Before & After
AI Simulations vs Clinical Before-and-After Photos
AI simulations and clinical galleries answer different questions. Learn what each can show, what neither can prove, and how to present both without confusing a visual preference with an outcome.
8 min read - IIAI Before & After
AI Before & After Photo Privacy: A Practice Checklist
A lifecycle checklist for permission, processing, storage, access, training rights, retention, deletion, analytics, and vendor review when an aesthetic AI product handles facial photos.
10 min read - IIIAI Before & After
How Med Spas Can Use AI Before & After in Consultations
A provider-led workflow for using an illustrative AI preview to clarify preferences, surface expectation gaps, and preserve the line between a visual aid and the actual treatment plan.
9 min read