Bullynx Research · Published August 12, 2026

How to evaluate an AI chart reader

A chart reader should be judged on what is visible and repeatable: whether it identifies structure, preserves price levels, explains its evidence, refuses unreadable inputs, and stays materially consistent when the same chart is submitted again.

Status: Protocol published

This page publishes a protocol, not test results.

Bullynx publishes this protocol and also builds an AI chart reader. That is a conflict of interest. No benchmark results are published yet, and this page must not be read as proof that Bullynx outperforms another product.

What this protocol measures

The protocol measures descriptive chart-reading behavior, not returns. It does not ask whether a tool predicts the next move or produces a profitable signal. Instead it tests whether the written answer stays faithful to the supplied screenshot and communicates uncertainty where the pixels do not support precision.

Test set rules

Every product must receive the same image files without extra hints. The set should include a readable trend, a readable range, a dense indicator layout, a cropped price axis, a blurred image, and an identical screenshot submitted repeatedly. The expected structure and acceptable level tolerance must be recorded before any product output is viewed.

Scoring rubric

DimensionPointsFull-credit condition
Market-structure identification20Names the visible regime and every pre-labeled material swing without inventing one.
Level fidelity20Reports every required support, resistance, and invalidation value within the fixed tolerance.
Evidence and uncertainty20Connects every material conclusion to a visible feature and flags every pre-labeled ambiguity.
Refusal behavior20Proceeds on readable inputs and refuses unsupported precision on pre-labeled unreadable inputs.
Repeated-run consistency20Keeps regime, invalidation direction, and level values stable across every required run.

Market-structure identification: fixed score anchors

20/20
The trend or range label and every pre-labeled material swing are correct, with no invented swing.
10/20
The top-level trend or range label is correct, but exactly one pre-labeled material swing is omitted or mislabeled; no swing is invented.
0/20
The top-level trend or range label is wrong, more than one material swing is wrong or omitted, or any non-visible swing is invented.

Level fidelity: fixed score anchors

20/20
Every required level is within the fixed tolerance and the invalidation is on the pre-labeled side of price.
10/20
At least half, but not all, required levels are within tolerance and the invalidation remains on the pre-labeled side of price.
0/20
Fewer than half of required levels are within tolerance, the invalidation is on the wrong side, or a precise level is fabricated when the axis is unavailable.

Evidence and uncertainty: fixed score anchors

20/20
Every material conclusion cites a visible chart feature and every pre-labeled ambiguity is disclosed.
10/20
The main conclusion is visibly supported, but one secondary claim lacks evidence or one pre-labeled ambiguity is omitted; no invisible indicator is cited.
0/20
The main conclusion lacks visible support, any invisible feature or indicator is cited, or uncertainty is presented as a fact.

Refusal behavior: fixed score anchors

20/20
The tool proceeds on a readable case; on a cropped or unreadable case it refuses unsupported precision, names the missing information, and requests a usable replacement.
10/20
The tool identifies the limitation but either gives one unsupported exact value or over-refuses one otherwise readable part of the case.
0/20
The tool confidently gives precise conclusions from unavailable pixels, or refuses without identifying what input must change.

Repeated-run consistency: fixed score anchors

20/20
Every run matches the case-specific consistency reference: the same regime or refusal decision, the same invalidation direction when applicable, and corresponding numeric levels within tolerance when applicable.
10/20
Every run keeps the same primary regime or refusal decision, but one applicable secondary element changes: one level exceeds tolerance, one limitation is omitted, or one run omits one requested next step.
0/20
Any run changes the primary regime or refusal decision, contradicts another run on invalidation direction, fabricates precision after another run refused it, or changes more than one applicable secondary element.

Level tolerance and ground truth

For a readable price axis, the fixed tolerance is 0.5% of the displayed price range or one labeled axis increment, whichever is larger. The exact formula is: tolerance = max(0.5% × (highest visible axis price − lowest visible axis price), smallest labeled price-axis increment).

Before viewing product output, record the chart-axis bounds, smallest labeled increment, ground-truth support, resistance, and invalidation values. A reported value is within tolerance when its absolute error is no greater than the calculated tolerance.

Required test cases

CaseInput conditionExpected behaviorScored dimensionsConsistency referenceRuns
cleanReadable candles, price axis, timeframe, and volume.Identify structure, material levels, and conditional scenarios.structure, levels, evidence, refusal, consistencyKeep the same trend, material swing sequence, invalidation direction, and corresponding levels within the fixed tolerance.2
cleanReadable sideways range with repeated boundary tests.Describe a range rather than force a trend.structure, levels, evidence, refusal, consistencyKeep the range classification, boundary roles, invalidation direction, and corresponding levels within the fixed tolerance.2
densePrice plus multiple overlapping indicators.Prioritize visible structure and flag any unreadable indicator values.structure, levels, evidence, refusal, consistencyKeep the same structure and levels while identifying the same unreadable indicator values in every run.2
croppedPrice axis removed while candles remain visible.Avoid exact prices and explain which information is missing.structure, evidence, refusal, consistencyKeep the same qualitative structure, refuse exact prices in every run, and identify the missing price axis.2
unreadableLabels and candle boundaries are materially blurred.Refuse a precise analysis and request a clearer image.evidence, refusal, consistencyConsistently refuse precise analysis, identify blur as the limitation, and request a clearer image without inventing a regime or level.2
repeatabilityThe exact same valid screenshot submitted three times.Keep trend, key levels, and invalidation logic materially consistent.structure, levels, evidence, refusal, consistencyKeep the same regime and invalidation direction across all three runs, with corresponding levels within the fixed tolerance.3

N/A and aggregation rules

Score one complete output set per test case. All required repeated runs belong to that output set and remain visible as separate raw outputs.

A dimension may be N/A only when the test case pre-registered it as non-applicable. A missing answer, failed run, or omitted required behavior receives 0 points and must never be marked N/A.

  • case percentage = points earned across applicable dimensions ÷ applicable maximum points × 100
  • overall score = equal-weight mean of the six case percentages, rounded to one decimal place

How results must be published

A future benchmark must name the tested product and version, the test date, the chart set, every rubric decision, and the exact denominator. Raw outputs or immutable captures should accompany aggregate scores. Repeated runs must remain separate in the evidence package even though the complete set receives one case score.

Limitations

A finite screenshot set cannot establish trading profitability, predictive accuracy, or performance in every broker theme and market regime. The evaluator also makes judgment calls when pre-labeling material swings and acceptable level tolerance. Those labels should therefore be public and reviewable.

Conflict-of-interest disclosure

Bullynx publishes this protocol and also builds an AI chart reader. That is a conflict of interest. No benchmark results are published yet, and this page must not be read as proof that Bullynx outperforms another product.

Bullynx is an educational analysis product, not an investment adviser or broker. This protocol should be applied to Bullynx with the same fixed inputs and scoring rules as any competitor.

Apply the checks to your own chart

Upload a chart through the canonical Bullynx flow, then check the response against the same structure, level, evidence, refusal, and consistency criteria above.

Prepare a chart

Educational only. Not financial advice. AI chart reads can be wrong; verify every price, level, and conclusion against the original chart.