Why Face Rating Apps Give You Different Scores
Five reasons the same photo gets different numbers from different face analysis apps, how to test any app yourself in ten minutes, and what a tool should be telling you.

Upload the same photo to three face analysis apps and you will get three different answers. Sometimes wildly different. Upload the same photo to the same app twice, taken minutes apart in different light, and you will often get two different answers from that.
This is not a mystery and it is not a sign that one app has secret insight the others lack. There are five specific reasons, they compound, and once you know them you can test any tool yourself in about ten minutes.
1. They are not measuring the same thing
The most common reason, and the least discussed.
Almost every facial measurement has more than one accepted convention, and apps rarely tell you which they use.
Canthal tilt can be measured against true horizontal or against the interpupillary line, the line running through both pupils. If your head is tilted three degrees in the photo, those two methods differ by three degrees. On a measurement whose entire population range is roughly eight degrees wide, that is an enormous discrepancy produced purely by convention.
Facial width-to-height ratio has two definitions in circulation. The one used in most of the research literature measures height from the upper eyelid to the upper lip. Another measures from the brow. The second produces a smaller number for the same face, every time.
Midface ratio depends on where you decide the midface starts and stops. Pupil line or nasion at the top. Upper lip or subnasale at the bottom. Four combinations, four different numbers, none of them wrong so much as differently defined.
Gonial angle measured on a photograph is measuring the soft tissue outline of your jaw. Measured on a lateral cephalogram in a clinic, it is measuring bone. These routinely disagree by ten degrees or more depending on how much tissue sits over the mandible.
Face shape depends on where the hairline reference point is placed, which on a receded hairline is a judgement call that can move you across a category boundary.
None of this is malpractice. It is a real ambiguity in the underlying measurement. The problem is that almost nobody discloses which convention they picked, which means you cannot compare two results even in principle.
2. Your photo is doing more work than your face
Measurement error starts at the camera, and most of it is invisible to you.
Camera distance. A phone held at typical selfie range is close enough that perspective distortion becomes substantial. The centre of the face projects larger relative to the periphery, so your nose reads wider and longer while your jaw and ears compress. Every width-based measurement is affected. The same face photographed from arm's length or further with a slight zoom produces measurably different numbers, and the further shot is closer to how you look in person.
Head rotation. Yaw, meaning turning left or right, compresses one side of the face. Pitch, meaning chin up or down, compresses or extends the vertical sections and directly changes facial thirds, midface ratio, and FWHR. Roll, meaning tilt, rotates the whole reference frame and corrupts any angle measured against image horizontal.
A few degrees of rotation you would not notice in a photo produces errors larger than the differences people are trying to detect.
Lighting. Directional light creates shadow on one side of the face. Shadow edges are high-contrast lines, and landmark detection models will happily latch onto a shadow boundary instead of the anatomical feature underneath. This is the single largest source of false asymmetry, and it is why the same person can look symmetric in one photo and noticeably asymmetric in the next.
Expression. Smiling raises the cheeks, changes the mouth landmarks, and narrows the eye aperture. Raised brows change every brow-referenced measurement. Most people unconsciously raise their brows slightly when being photographed.
Focal length. Different phones and different camera modes use different focal lengths, which changes the degree of perspective compression independently of how far away you held it.
If a tool does not check for any of this before measuring, it is measuring your photo conditions and reporting them as facts about your face.
3. Landmarks are not points
This one is structural and no tool escapes it entirely.
Facial measurement works by locating specific anatomical landmarks and computing distances and angles between them. Some landmarks are genuinely well defined. The centre of the pupil is a point.
Many are not.
The lateral canthus, the outer corner of the eye, is a soft junction where the lids meet. Its apparent position moves with lighting, squinting, and camera angle.
The gonion, the corner of the jaw, is not a corner at all on a soft tissue outline. It is a curve, and where you place the vertex changes the angle you compute.
The zygion, the widest point of the cheekbone, is a maximum along a smooth arc, and soft tissue displaces it.
The trichion, the centre of the hairline, does not exist in any stable sense on a receded hairline.
Different landmark models place these differently, because there is no single correct answer. Two well-built tools using two well-trained models will disagree here, and both will be reasonable.
The difference between tools is whether they tell you. A tool that reports a gonial angle to one decimal place, from a landmark it inferred from a curve, is presenting a precision it does not have.
4. Composite scores hide their own arithmetic
If an app gives you a single number out of ten, that number was produced by combining individual measurements according to weights.
The weights are the entire product. They determine whether your canthal tilt matters more than your facial thirds, whether symmetry counts for five percent or thirty, and what "good" means for each input.
Those weights are essentially never published.
This means two things. Two apps with identical measurements can give you scores several points apart purely because they weighted things differently. And you cannot evaluate whether the weighting is sensible, because you cannot see it.
A composite attractiveness score is an aggregation choice presented as a finding. The choice is doing all the work, and the choice is private.
5. Some apps are not measuring at all
There is a category difference that matters more than any of the above.
Deterministic geometry means the app locates landmarks and computes measurements with explicit formulas. Given landmark coordinates and the formula, you could reproduce the number by hand. It can be checked, explained, and corrected.
Learned scoring means the app passes your photo through a model trained to predict human attractiveness ratings, and outputs whatever the model says. There is no formula. There is no landmark you can point to. When it is wrong, nobody can say why.
Learned scoring has an additional problem that is rarely acknowledged. A model trained to predict ratings learns the preferences of whoever produced the training ratings. If those raters were demographically narrow, the model has encoded that narrowness, and it will apply it to everyone. The model does not know it is doing this and neither does the app.
Both approaches exist in this market and they mostly look the same from the outside. Both give you a number.
Why the same app gives you different answers
Not just between apps. Within one.
Photo conditions, per everything above. Different light, different distance, different head angle, different result.
Periorbital swelling. Your eye area genuinely differs day to day based on sleep, sodium, alcohol, and allergies. That is a real change in your face, not a measurement error.
Body composition over time. Facial fat sits over the jaw and cheekbones. This one is a real change too.
Model updates. Apps update their models without announcing it. A number from six months ago may not have been produced by the same method as today's number.
Nondeterminism. Some pipelines involve stochastic steps and do not return the same answer for the same input twice.
That last one is testable, which brings us to the useful part.
How to test any face analysis app in ten minutes
Four tests. Any tool that fails these is telling you less than it appears to.
Test 1: The same photo twice
Upload one photo. Note the numbers. Upload the identical file again.
Should happen: identical results. If they differ: the pipeline is nondeterministic, which means at least part of the variation you see between sessions is noise rather than your face.
Test 2: Rotate slightly
Take two photos in the same light and position, one facing straight on, one with your head tilted maybe five degrees to the side. Not enough to look obviously tilted.
Should happen: either near-identical angular measurements, or a rejection asking you to retake. If canthal tilt shifts by roughly the amount you tilted: the tool is measuring against image horizontal rather than the interpupillary line, and your result is partly a measure of how you held your head.
Test 3: Change the lighting
Two photos, same position, one in flat even light, one with a lamp or window on one side.
Should happen: symmetry results should be close, or the tool should flag the directional light. If symmetry changes noticeably: it is measuring shadow.
Test 4: Ask it how it knows
Look for a methodology page. Look for the landmark definitions. Look for which reference line it used. Look for a confidence score.
If none of that exists, the tool has made a series of consequential choices on your behalf and declined to tell you what they were.
What a tool should be telling you
Not a request. A checklist.
Which convention it used. Reference line, landmark definitions, measurement method. Per metric.
How confident it is, per measurement. Not one overall confidence. Some measurements are structurally harder than others, and a good tool says which.
Why confidence is low when it is. "Head rotation detected" is useful. A bare percentage is not.
When it cannot measure something. Occluded landmarks should produce "unavailable," not an estimate presented as a result.
Ranges rather than ideals. Target figures circulating for these measurements mostly trace back to small aesthetic preference studies or to classical artistic convention, not to population data. A tool reporting your distance from an ideal is reporting distance from an assumption.
What it cannot do. Whether a single frontal photo can support the measurement at all. Whether the number is measuring bone or soft tissue. Where the error bars actually sit.
The honest summary
Facial measurement from a photograph works, within limits, if the conditions are controlled and the method is stated. Ratios are more reliable than absolute distances. Well-defined landmarks are more reliable than inferred ones. Frontal measurements are more reliable than anything requiring depth.
What does not work is a single number telling you how attractive you are. Roughly half of attractiveness judgement comes from the individual taste of the person looking, which no measurement predicts. The other half has a real shared component, but resolving it to one decimal place from one photograph is not something anyone can currently do.
If the numbers you are getting keep changing, the most likely explanation is that they were never as precise as they were presented.
How we measure, in full · Measure your face
Frequently asked questions
Are face rating apps accurate? For proportional measurements under controlled photo conditions, reasonably so, within stated error. For attractiveness scores, no. A score is a weighted combination of measurements, and the weights are a private choice rather than a finding.
Why do I get different scores on different apps? Different measurement conventions, different landmark models, different photo quality handling, and different scoring weights. Any one of these can produce a large difference by itself.
Why does my score change on the same app? Photo conditions are the usual cause, particularly lighting, camera distance, and head angle. Real day-to-day changes in periorbital swelling also contribute. Some pipelines are also nondeterministic and simply do not return identical results for identical inputs.
Which face analysis app is most accurate? The useful question is which one tells you its method. An app that publishes its measurement conventions, landmark definitions, and confidence derivation can be checked. One that outputs a score with no stated method cannot be evaluated at all, by you or by anyone.
Does camera quality affect the results? Distance and focal length matter considerably more than sensor quality. A photo taken close to the face distorts proportions substantially regardless of how good the camera is.
Should I trust a face rating out of ten? Treat it as a description of that specific photograph filtered through an undisclosed weighting. Individual proportional measurements, with their methods stated, are more informative and more honest.
Related reading How ASCNDED measures a face · Full methodology · How attractive am I? · The PSL scale explained
See your own numbers.
Upload one photo and get your facial proportions, face shape, and a non-surgical plan in about 60 seconds.
Scan my face