Platform Why Features Security Score AI Engine AI Coding KYP Hub Pricing Company About Buckler News Contact Français Book Demo →
Study Design

Methodology

What we expected going in, how the test was built to check it, and exactly how each trial was run.

1
Hypothesis
What we expected commercial AI platforms to get wrong, and why

Commercial AI tools, because they lack access to fundamental or primary data, are unable to answer basic research questions that require pulling and comparing information across many securities at once. They are well suited to working on a single security in isolation — reading an annual report, summarizing a fund's stated strategy, or working from a specific filing that functions as a primary source in its own right. They are not equipped to pull verified, current performance data for hundreds of mutual funds, ETFs, or equities and rank them against one another.

The Question

  • Can a commercial AI platform correctly identify the top five Canadian mutual funds by three-year return?
  • Can it correctly identify the top five Canadian-listed ETFs by three-year return?
  • Can it correctly identify the top five Canadian equities by trailing return?

Hypothesis: No. Because none of these platforms have access to a complete, current, and structured dataset of fund, ETF, and equity performance, no commercial AI platform will be able to produce a correct, verifiable answer to any of the three questions above.

The reason isn't carelessness. It's how these platforms are built. An AI engine doesn't query a database and sort the results — it generates a response by pattern-matching across the vast body of financial text it was trained on, producing whatever fund names, tickers, and return figures are statistically likely to sound correct based on what it has seen discussed most often. That output is then formatted exactly like a real answer: a clean table, specific percentages, and often a cited source. Nothing about the presentation signals whether a figure was retrieved from a live source or reconstructed from memory — the two look identical on the page. Because the model is trained to produce a complete, confident-sounding response rather than flag what it doesn't know, it ends up naming the securities most discussed in its training data — not the best performers — and attaching real-looking numbers to them with no indication that the underlying process was recall, not retrieval.

2
Design
How the test was built to confirm or disprove the hypothesis

The design goal was replicability: any advisor with access to the same six platforms should be able to run this test themselves and land on comparable results. That meant fixing every variable we could control — which platforms, which account tier, which prompts, in which order, over how many sessions — so that whatever differences showed up in the output were attributable to the engines, not to how the questions were asked. The exact platforms, versions, and session parameters used are reproduced in full under Replication Parameters in the Appendix.

The study measured three things against each engine: consistency (did it give the same answer to the same question asked three separate times), accuracy (did its answer match verified, current market data), and transparency (would it disclose, when asked directly, where its numbers actually came from and how confident it was in them). A platform could fail on any one of these independently of the others — an engine could be consistently wrong, or accurately right by chance in one trial and wrong in the next, or simply refuse to answer the transparency questions honestly.

Six commonly available commercial AI platforms were tested, each on its paid or pro consumer tier, using the standard consumer chat interface rather than a developer API, an agent mode, or a custom GPT/persona. This was a deliberate choice: it reflects how a financial advisor would actually use these tools day to day. The exact model versions, retrieval tooling, and knowledge cutoffs used for each — ChatGPT, Claude, Copilot, Gemini, Grok, and Perplexity — are listed in the Appendix.

Five prompts were used, and the same five were used for every engine, worded identically every time. Prompts 1, 2, and 3 are the core performance questions — one per asset class — and each was run as the opening prompt of its own session. Prompts 4 and 5 are diagnostic: they don't ask for new performance data, they interrogate the answer the engine just gave, immediately afterward, in the same session. Each engine ran each asset-class chain three times, in a fresh private/incognito session every time, so nothing carried over from memory between trials.

Asked each engine for the top five Canadian equities by trailing 12-month return, filtered to Canadian-domiciled issuers with market cap over $1 billion CAD, with six supporting data points per security — market cap, share price, P/E, dividend yield, 3-year price appreciation, and beta.

Asked each engine for the top five Canadian equity or Canadian-focused equity mutual funds by 3-year annualized return, with nine supporting data points per fund — FundServ code, MER, 3- and 5-year return, standard deviation, drawdown, AUM, and CIFSC category.

Asked each engine for the top five Canadian-listed ETFs by 3-year annualized return, with seven supporting data points per ETF — ticker, issuer, MER, AUM, 3-year and 1-year return, and average daily volume.

Run immediately after every core performance prompt, this asked the engine to account for its own answer: enumerate every source it relied on, classify each as primary or context-only, disclose whether each was actually fetched during the session or only recalled, rate its own confidence per data point, and state directly whether it had reached Bloomberg, FactSet, Morningstar Direct, S&P Capital IQ, or SEDAR+. This is the prompt behind the sourcing findings discussed in Citations & Sourcing below.

Closed every session by asking the engine to document its own technical conditions — model version, knowledge cutoff, whether live retrieval was active, and similar metadata — so a later replication attempt has something concrete to compare its own session against.

The verbatim wording of all five prompts, the complete list of required data points and disclosure fields for each, and the conditions under which they were run are reproduced exactly under Replication Parameters in the Appendix.

3
Protocol
How every trial was run, and how the resulting 54 trials were assembled

Every individual trial ran in its own private, incognito browser session, opened immediately before the trial began and closed immediately after the response was captured. An AI engine can carry context from one query to the next within an open session, so starting fresh for every single trial — not just every asset class, not just every engine — was the only way to guarantee that no trial was quietly informed by anything that came before it.

Within a given engine, the three asset classes were run back-to-back: one asset class was taken through all three of its trials before moving to the next, each trial using the same three-prompt chain for its asset class (the core performance prompt, followed by the Prompt 4 sourcing audit, followed by the Prompt 5 session-metadata capture). The same sequence was then repeated, asset class by asset class, for each of the six engines, producing 54 trials in total. A worked example of the full nine-trial sequence for one engine, and the complete execution count across all six, are reproduced under Replication Parameters in the Appendix.

Once all 54 responses were captured, each was coded along the same three dimensions the study was designed to measure — consistency, accuracy, and transparency, defined precisely under the Appendix's Scoring Method. That coding is what produced the dispersion map in the Findings section below, and the verified-versus-returned comparisons in the Accuracy section that follows it.

Results

Findings

Across 54 trials, three findings repeated themselves regardless of engine, asset class, or trial number. No engine consistently produced the same answer twice. No engine returned a verified top performer in any category. And no engine disclosed the limits of its data unless directly asked.

1
Consistency
Where engines agreed with themselves, they were consistently wrong

Consistency was tested by asking each engine the identical question three separate times, in three fresh sessions, and comparing what came back. Across 54 trials the results split into two failure patterns, and neither one is reassuring. Engines were either locked onto the same wrong answer every time, or they were volatile — swinging between unrelated answers to the exact same prompt, with no warning in the output that anything had changed.

Where engines were consistent, they were consistently wrong. Grok produced the same five ETFs in identical order across all three ETF trials — the highest ETF consistency of any engine in the study, and the audit prompt confirmed it reflected a memorised list rather than any live screening. Perplexity's ETF Trials 2 and 3 were byte-identical, consistent with the same pattern: a stable training-data snapshot standing in for a real answer. Claude's top four ETF picks held constant across all three ETF trials, making it the most internally stable engine on that asset class — and Claude's third mutual-fund trial was refused outright, a refusal that reproduced identically on a rerun. Even a refusal, in other words, can be a "consistent" result. An engine that gives you the same wrong answer, or the same non-answer, three times in a row is not consistent in any meaningful sense. It is stuck.

The more common and more troubling pattern was significant variation between trials of the identical prompt. Gemini posted the worst record in the study on both fronts at once: zero live retrieval in all nine of its trials, fabricated as-of dates on every figure, and the lowest cross-trial consistency of any engine — it did not even reliably repeat its own fabrications. Copilot's equity Trials 1 and 2 were byte-identical before Trial 3 returned a completely unrelated list of micro-cap stocks with no overlap from the previous two — and in a separate Equity Trial 2 run, Copilot's own Prompt 4 disclosure described its own numbers as "FABRICATED PLACEHOLDERS, not retrieved from any source," with nothing in the original table hinting at that. Perplexity refused two of its three equity trials outright and answered the third, meaning the same prompt, asked three times, produced a coin-flip between a refusal and a full formatted answer.

Same Prompt, Three Different Answers — ChatGPT, Top-Five ETFs

ChatGPT was asked the identical top-five Canadian ETF question three times, in three fresh incognito sessions, using the exact same wording each time. Each trial came back with an entirely different category of fund, and not one fund carried over from any trial to any other:

  • Trial 1: Leveraged gold ETFs
  • Trial 2: All-bitcoin ETFs
  • Trial 3: Broad gold-mining equity funds

Same engine. Same prompt. Same account tier. Three trials in a row, three unrelated investment theses — precious-metals leverage, cryptocurrency, and mining equities — each presented with the same confident formatting as the last. Nothing in any single response would tell a reader that the other two trials existed, let alone that they disagreed entirely.

Celestica appeared in 14 of 18 equity trials — the single most repeated pick across every engine and every trial. Its actual one-year return placed it outside the verified top five by more than 300 percentage points. That level of consensus tells us nothing about performance. It tells us which company generated the most financial media coverage, analyst attention, and training-data presence. Agreement between engines, like agreement between an engine's own trials, turned out to be a measure of shared exposure to the same training and search content — not a measure of being right. The full citation count behind every figure in this section — every security, every engine, every trial — is reproduced in the Appendix.

Dispersion Map — How Each Engine Sourced Its Answers
P
Pulled from a real source
S
Used search snippets
I
Guessed from training data
M
Made it up from memory
F
Fabricated source citations
–
Refused / no answer
AI Engine Equities Mutual Funds ETFs
T1T2T3 T1T2T3 T1T2T3
ChatGPTSIISIISPI
ClaudeSSSSS–PPS
CopilotSMSSMMSMM
GeminiMMMMMMMMM
GrokFIIIIIIIF
Perplexity–SSSSSSIS
2
Accuracy
The gap between what the engines returned and what the verified data showed

Accuracy was checked the only way that means anything: against Buckler's verified market data for the same period, as of April 30, 2026. Across all three asset classes and all 54 trials, the gap between what the engines returned and what the verified data showed was not marginal. It was categorical. The engines were not slightly off on the numbers for the right securities. They were identifying a different set of securities entirely — ones that happened to be well-known, widely covered, and prominent in financial media, regardless of their actual performance by the metric each prompt specified.

Citation Rate Legend

Every "Citation Rate" figure below is the share of that asset class's 18 trials (six engines, three trials each) in which the engine's answer included that specific security.

14% Under 20% of trials
50% 20%–80% of trials
86% Over 80% of trials

The verified top five Canadian equities by one-year trailing return were Almonty Industries (+685.9%), Groupe Dynamite (+587.5%), Hut 8 Mining (+506.8%), Faraday Copper (+460.5%), and Spartan Delta (+417.3%). Across all 18 equity trials — six engines, three trials each — four of those five names were never cited once. Faraday Copper was cited exactly once, a 5.6% citation rate. Almonty Industries, the single best-performing Canadian equity in the entire dataset, up nearly 686% over the period, was named by zero engines in zero trials: a 0% citation rate.

CompanyTickerCitation Rate1-Year Return
Almonty Industries Inc.AII0.0%+685.9%
Groupe Dynamite Inc.GRGD0.0%+587.5%
Hut 8 Mining CorpHUT0.0%+506.8%
Faraday Copper CorpFDY5.6%+460.5%
Spartan Delta CorpSDE0.0%+417.3%
Celestica Inc. (most-cited pick)CLS77.8%+374.6%
Aritzia Inc.ATZ50.0%+195.9%
Bombardier Inc. (Class B)BBD.B38.9%+257.4%

Celestica was the engines' consensus pick — a 77.8% citation rate, more than any other name by a wide margin — and its actual return placed it more than 310 percentage points below the verified number one. That is the clearest illustration in the study of what "consensus" from these engines actually measures: not performance, but how often a name shows up in the training data and search results the engines draw on. Celestica is a large, widely covered technology manufacturer. Almonty Industries, a smaller mining company, is not — and that difference in media visibility, not in return, is what determined which one the engines converged on.

The verified top five Canadian equity mutual funds by three-year annualized return were Purpose Global Resource Fund Series L (+52.3%), Friedberg Global-Macro Hedge Fund U$ (+50.1%), CI Precious Metals Fund Series I (+50.1%), Dynamic Precious Metals Fund Series O (+48.3%), and Ninepoint Silver Equities Fund Series D (+47.3%). None of the five appeared in a single trial, by any engine, at any point in the study — a 0% citation rate across all 18 mutual fund trials.

FundCitation Rate3-Year Return
Purpose Global Resource Fund Series L0.0%+52.3%
Friedberg Global-Macro Hedge Fund U$0.0%+50.1%
CI Precious Metals Fund Series I0.0%+50.1%
Dynamic Precious Metals Fund Series O0.0%+48.3%
Ninepoint Silver Equities Fund Series D0.0%+47.3%
Guardian Canadian Focused Equity Series F (most-cited pick)16.7%+23.8%
RBC Canadian Equity Fund Series F27.8%+19.4%

The fund the engines cited most often, Guardian Canadian Focused Equity Series F, still only reached a 16.7% citation rate — under the 20% line — and returned 23.8% over three years, less than half of the actual top performer. And the miss went beyond picking the wrong fund: four of the 18 mutual fund trials didn't return mutual funds at all. The engines substituted ETFs instead, meaning in more than one in five trials, the response failed to match even the asset class the prompt specified, before any question of which specific fund was correct.

The Real Top Ten, Named Once

Combine the verified top five Canadian equities with the verified top five Canadian equity mutual funds — ten securities, each the actual best performer in its category over the relevant period. Across the 36 trials that tested those two asset classes (18 equity trials and 18 mutual fund trials, six engines run three times each), those ten securities were named by an AI engine a combined total of once: Faraday Copper, in a single equity trial.

The other nine — including Almonty Industries, the single best-performing security in the entire study at +685.9% — were never named by any engine, in any trial, at any point in the research. Not once, across six platforms and eighteen independent attempts per asset class.

The verified top five Canadian-listed ETFs by three-year annualized return were concentrated in gold and precious metals: BMO Equal Weight Global Gold Index ETF and iShares S&P/TSX Global Gold Index ETF (tied at +53.7%), BetaPro Gold Bullion 2x Daily Bull ETF (+50.0%), BMO Junior Gold Index ETF (+48.7%), and Global X Global Semiconductor Index ETF (+47.5%). Unlike equities and mutual funds, the engines did land in the right neighbourhood here — gold ETFs were a recurring theme across trials — but never on the specific, correctly ranked winners, and the study's single most-cited ETF overall did not even meet the prompt's filter criteria.

ETFTickerCitation Rate3-Year Return
BMO Equal Weight Global Gold Index ETFZGD27.8%+53.7%
iShares S&P/TSX Global Gold Index ETFXGD38.9%+53.7%
BetaPro Gold Bullion 2x Daily Bull ETFGLDU5.6%+50.0%
BMO Junior Gold Index ETFZJG27.8%+48.7%
Global X Global Semiconductor Index ETFCHPS5.6%+47.5%
CI Galaxy Bitcoin ETF (most-cited overall — fails the prompt's filter)BTCX.B55.6%+36.7%

The engines' gold ETF appearances reflect training-data familiarity with a well-known category of fund, not a live performance screen that correctly ranked one gold ETF over another — none of the five verified winners cleared a 40% citation rate, and the semiconductor ETF that placed fifth sat at just 5.6%. More striking is what the engines cited most often overall: CI Galaxy Bitcoin ETF, at a 55.6% citation rate, ahead of every verified top-five gold ETF. It does not meet the Canadian-equity filter specified in the prompt, and its 36.7% three-year return sits nearly 17 points below the actual top performer. The engines' most confident, most frequent answer was not just short of the correct one — it was disqualified by the prompt's own criteria.

Every citation count and verified return referenced in this section — including every security cited but not shown above — is reproduced in full in the Appendix, as a standing reference for checking any accuracy claim in this document against the underlying data.

3
Citations & Sourcing
What the engines cited, what they actually retrieved, and what they admitted when pressed

Prompt 4, run immediately after every core performance question, asked each engine to account for itself directly: where did you source this information? Where did most of the sourcing actually come from? How did you obtain the data? And, specifically, where did you actually get it, as opposed to where you said you got it? The answers, run across all 54 trials, establish something the original responses never disclosed on their own.

None of the six engines accessed any institutional financial database in any trial. Bloomberg, FactSet, Morningstar Direct, S&P Capital IQ, SEDAR+, and exchange-level data feeds — the tools that would actually support a ranked, verifiable screen across hundreds of securities — were never reached, not once, by any engine. Several engines named these platforms in their output anyway. What the engines used in practice was a mix of training data, web search snippets, and in several cases, nothing at all.

An AI engine is not a search engine and it is not a database. It is a language model trained to predict what a helpful, coherent response looks like. When it cannot retrieve the data a question requires, it does not stop. It produces what a correct answer would look like based on patterns in its training data — including plausible figures, plausible fund names, and plausible citations. Every engine in this study has a knowledge cutoff of between late 2023 and mid-2024, leaving a gap of twelve to twenty months between its last training snapshot and the May 2026 prompt. A response that includes a formatted table, specific return figures, named sources, and an as-of date looks like a researched answer. There is no way to tell from the output alone that it is not — the only way to know is to ask, and the engines were asked in every one of the 54 trials.

The clearest illustration of the gap between a claimed source and an actual one came from Perplexity, the engine that cited more paywalled institutional sources than any other in the study — ten mentions across all trials — without retrieving a single one of them. In one mutual fund trial, the engine's own output referenced Morningstar Direct, an institutional research terminal that costs thousands of dollars per seat annually and implies direct, licensed data access. When asked under Prompt 4 to confirm whether it had actually reached Morningstar's terminal, the audit told a different story:

"Morningstar Direct in prior answer was QUOTING Morningstar's own article description, NOT terminal access."Perplexity · Mutual Fund Trial 1 · Prompt 4

The distinction is not a technicality. One is a data retrieval, priced and licensed specifically because it is authoritative. The other is a name appearing in the first sentence of a public webpage, reproduced as if it were a source attribution. A reader relying on the original response — the one without the audit prompt attached — has no way to tell the two apart, because the engine formatted its citation of the article description identically to how it would have formatted a genuine terminal pull.

Grok went a step further: in two of its nine trials, it produced source citations that were formatted correctly, named plausible-sounding institutions, and appeared directly alongside figures presented as retrieved data. There was nothing about the citations themselves that would flag them as suspect. When audited, Grok acknowledged the figures had not come from any specific source at all:

"Approximated from snippets and general knowledge of commodity ETF behavior during a strong metals period."Grok · ETF Trial 1 · Prompt 4

A citation that names a real institution but was never actually consulted, and a citation that was fabricated outright and attached to an approximated figure, produce the same output on the page: a footnote that looks sourced. The only way to tell the difference between a genuine citation, a misattributed one, and an invented one is to ask the engine to account for it directly — and every engine in this study required that follow-up before it would say so.

"FABRICATED PLACEHOLDERS, not retrieved from any source."Copilot · Equity Trial 2 · Prompt 4
"The specific numeric values were simulated projections rather than retrieved market data."Gemini · Equity Trial 1 · Prompt 4
"Ranking response should NOT be treated as a compliance-grade investment screening, an institutional due-diligence report, or an audit-ready performance ranking."ChatGPT · ETF Trial 2 · Prompt 4
"The 1.14% MER I assigned to the Mawer Canadian Equity Fund (Series A) was likely pulled from the Dynamic Canadian Dividend Fund description in the Wealth Professional snippet. These are different funds. This was an error."Claude · Mutual Fund Trial 1 · Prompt 4

Run across all six engines, Prompt 4's four questions — where the information was sourced, where most of the sourcing actually came from, how the data was obtained, and where it actually came from versus where the engine said it came from — produced a consistent pattern: what the original response implied and what the audit revealed were rarely the same thing.

EngineWhat the Response ImpliedWhat Prompt 4 Actually RevealedInstitutional Databases Reached
ChatGPTA researched, formatted ranking with figures and an as-of date4 of 9 trials from search snippets, 4 of 9 from training data alone, 1 primary fetch — an ETF issuer page for a bitcoin product that failed the prompt's own filterNone
ClaudeSourced, dated figures with named references3 primary-source PDFs fetched (ETF trials only); elsewhere, 18 HTTP-403 errors, 8 permission errors, 7 refusals, and 8 JavaScript-empty fetches — the most access failures disclosed by any engineNone
CopilotA standard, complete results tableZero live retrieval in 6 of 9 trials; one trial's own audit response called the output "FABRICATED PLACEHOLDERS, not retrieved from any source"None
GeminiCurrent figures labelled "as of May 2026," despite deep integration with Google SearchZero live retrieval in 9 of 9 trials, confirmed on audit; every figure was a training-data projection presented as currentNone
GrokA results table with named source citationsFabricated citations in 2 of 9 trials; figures "approximated from snippets and general knowledge," not retrieved from any named sourceNone
PerplexityCitations referencing Morningstar Direct and other paywalled institutional sources — 10 mentions across the studyConfirmed on audit that a Morningstar Direct citation was a public article description, not terminal access; none of the 10 paywalled sources cited were ever retrievedNone

The complete citation record behind these findings — every security any engine named, how many times, and against the verified return — is reproduced in full in the Appendix, as a standing reference for checking any consistency, accuracy, or sourcing claim made in this document against the underlying data.

Implications

Conclusions

A neatly formatted top-five table with citations and an as-of date can be entirely fabricated. The presence of a citation is not evidence that it was fetched.

1
What This Means
Four conclusions about what commercial AI can — and cannot — do for investment research

Across 54 trials, three failure patterns kept repeating regardless of engine, asset class, or trial number: no engine consistently reproduced its own answer, no engine returned a verified top performer in any category, and no engine disclosed the limits of its data unless asked directly. That isn't a list of bugs to patch. It's a structural limitation in what commercial AI can currently do with a research question that requires pulling and ranking real data across many securities at once. Four conclusions follow.

A neatly formatted top-five table with tickers, exact percentages, citations, and an as-of date looks identical whether every figure was retrieved from a live source or invented outright. The presence of a citation is not evidence that it was fetched — the Prompt 4 sourcing audit caught engines admitting, after the fact, that specific numbers were "simulated," "approximated," or "FABRICATED PLACEHOLDERS, not retrieved from any source." Nothing in the original response distinguished those figures from the ones that were real. An advisor reading the output has no way to tell which is which without a separate verification step the AI will not prompt them to take.

Where engines agreed — with each other, or with their own earlier trials — that agreement tracked media coverage and training-data presence, not returns. Celestica appeared in 14 of 18 equity trials, more than any other security in the study, and its actual one-year return placed it outside the verified top five by more than 300 percentage points. Repeating a question, or checking a second engine's answer against the first, does not independently confirm anything if both are drawing on the same underlying exposure. Cross-checking AI engines against each other is not verification.

This study used long, explicit, tightly filtered prompts — market-cap thresholds, CIFSC categories, exact ranking criteria, mandatory data points, "N/A" instead of a decline to answer — and it made no measurable difference. No amount of prompt engineering closes this gap, because the limitation isn't how the question is asked. It's that none of the six engines have access to institutional-grade primary data for Canadian securities: no Bloomberg terminal, no FactSet feed, no Morningstar Direct workstation, no FundServ record.

Closing that gap isn't simply a matter of an AI lab deciding to add it. Bloomberg, Morningstar, FactSet, and Thomson Reuters built their businesses on licensing that data at a metered, per-seat cost to institutions and professionals — often tens of thousands of dollars a year, priced around scarcity and exclusivity. Opening that data to a consumer AI platform would mean putting it, in practice, in front of every user of that platform — hundreds of millions of people — at effectively no marginal cost. That isn't a pricing adjustment for a data provider; it's the collapse of the business model the data was built to fund. There is no obvious commercial path by which Bloomberg-grade data becomes ChatGPT-grade data, because doing so would undercut the reason Bloomberg-grade data is valuable in the first place. That's a structural reason to expect this gap to persist even as the underlying models keep improving.

None of this is an argument against using AI. These engines are genuinely useful for orientation, idea generation, drafting, summarising, and brainstorming. They are not reliable for the kind of factual retrieval an advisor or investor needs without an explicit verification step — and the responses give no indication, on their own, that such a step is necessary. An advisor relying on this output risks presenting clients with securities selected by coverage rather than by returns, formatted with exactly the same confidence as if the numbers were real.

2
The Fix
What an AI engine would actually need to return replicable, consistent, and accurate results

The root cause is simpler than the failure modes make it sound: AI engines do not have access to high-quality primary data for Canadian securities, so when asked for a current return, MER, or AUM, they substitute whatever they can reach — rarely the right source, never a complete dataset. Better prompts will not fix a missing dataset. Producing replicable, consistent, and accurate results takes four things working together, not just a better-trained model.

Not web snippets, not training-data approximations, not a rate-limited free API tier — the same class of structured, current, institutional-grade data that Bloomberg, FactSet, Morningstar Direct, S&P Capital IQ, and SEDAR+ already provide to professionals. That access has to be licensed and paid for like any other institutional data feed; there is no free version of this that holds up.

The model itself should never be the source of a number. The figures a user sees need to come from a live query against a structured database — ranked and filtered deterministically — with the language model used only to interpret the question and format the answer, not to generate the data inside it. That is the difference between retrieval and recall, and it is the difference this entire study measured.

Every number displayed needs a traceable source, a retrieval timestamp, and a real as-of date attached automatically — not produced only when an engine is directly asked to disclose its sourcing, as the Prompt 4 audit had to force here. If a data point cannot be sourced, the honest answer is "not available," not a plausible-looking placeholder.

The same question, asked twice, needs to return the same answer, because it is drawing from the same underlying dataset rather than regenerating a plausible-sounding response from memory each time. Consistency in this study was mostly coincidence — engines that agreed with themselves were often just replaying the same memorised list. Real consistency comes from a shared source of truth, not a shared training set.

There is no shortcut between a consumer AI engine and a reliable investment research tool. The combination of real market data and purpose-built AI is not a premium option. For any firm that needs research it can stand behind, it is the only option.
Supporting Data

Appendix

The complete citation ledger each engine was measured against, the trial-by-trial breakdown behind it, and the parameters needed to replicate the study.

1
Full Results — Every Security Cited
The complete citation record behind the Consistency, Accuracy, and Citations & Sourcing findings above

The tables below list every security named by any engine, in any trial, across the study — the underlying ledger that the Consistency, Accuracy, and Citations & Sourcing findings are drawn from. Times Cited counts appearances across all 18 trials for that asset class (six engines, three trials each). Bolded, highlighted rows are the actual top five by verified return; as the Findings section shows, none were identified by any engine.

CompanyTickerTimes Cited1-Year Return
Almonty Industries Inc.AII0+685.9%
Groupe Dynamite Inc.GRGD0+587.5%
Hut 8 Mining CorpHUT0+506.8%
Faraday Copper CorpFDY1+460.5%
Spartan Delta CorpSDE0+417.3%
Celestica Inc.CLS14+374.6%
Energy Fuels Inc.EFR1+342.9%
Enerflex Ltd.EFX1+308.8%
Bombardier Inc. (Class B)BBD.B7+257.4%
Hammond Power SolutionsHPS.A4+230.5%
Kraken Robotics Inc.PNG1+228.6%
Aecon Group Inc.ARE1+218.7%
Aritzia Inc.ATZ9+195.9%
Cameco CorporationCCO5+169.0%
Finning International Inc.FTT3+160.7%
Sprott Inc.SII1+149.3%
IAMGOLD CorporationIMG2+134.1%
Kinross Gold CorporationK2+103.6%
Imperial Oil Ltd.IMO1+100.5%
Suncor Energy Inc.SU1+98.9%
TFI International Inc.TFII4+76.8%
Lundin Gold Inc.LUG5+71.3%
Agnico Eagle Mines Ltd.AEM3+59.3%
MDA Ltd.MDA1+54.4%
Wheaton Precious Metals Corp.WPM2+50.1%
Alamos Gold Inc.AGI2+38.0%
Shopify Inc.SHOP2+25.8%
TerraVest Industries Inc.TVK1−3.1%
Hemisphere Energy Corp. (below $1B cap — fails prompt filter)HME1—
Titan Mining Corp (below $1B cap — fails prompt filter)TI1—
HG Critical Capital (not confirmed on a major Canadian exchange)HG1—
Fossil Mining Corp (could not be confirmed to exist)—1—
FundTimes Cited3-Year Return
Purpose Global Resource Fund Series L0+52.3%
Friedberg Global-Macro Hedge Fund U$0+50.1%
CI Precious Metals Fund Series I0+50.1%
Dynamic Precious Metals Fund Series O0+48.3%
Ninepoint Silver Equities Fund Series D0+47.3%
Fidelity Canadian Growth Fund1+31.9%
Fidelity Special Situations Fund1+30.3%
TD Canadian Small-Cap Equity Fund2+27.2%
CI Morningstar Canada Momentum Index ETF (ETF returned as a fund)1+27.2%
Dynamic Power Canadian Growth Fund1+25.0%
DFA Canadian Vector Equity Fund4+24.6%
CI Morningstar Canada Value Index ETF (ETF returned as a fund)1+24.6%
Guardian Canadian Focused Equity Fund Series F3+23.8%
NCM Core Canadian Fund Series Z1+23.8%
SEI Canadian Equity Fund2+22.7%
Invesco RAFI Canadian Index ETF (ETF returned as a fund)1+22.6%
Guardian Canadian Focused Equity Fund Series A4+22.5%
Scotia Canadian Growth Fund Series A4+21.9%
DFA Canadian Core Equity Fund Series A1+21.3%
CI Select Canadian Equity Fund Series F1+20.9%
Vanguard FTSE Canada All Cap Index ETF (ETF returned as a fund)2+20.5%
Vanguard FTSE Canada ETF (ETF returned as a fund)1+20.5%
BMO S&P/TSX Capped Composite Index ETF (ETF returned as a fund)2+20.4%
iShares Core S&P/TSX Capped Composite ETF (ETF returned as a fund)2+20.4%
RBC QUBE Canadian Equity Fund3+20.0%
RBC Canadian Equity Fund Series F5+19.4%
iShares S&P/TSX 60 Index ETF (ETF returned as a fund)1+19.1%
Capital Group Cdn Focused Equity Fund Series F3+18.6%
Capital Group Canadian Focused Equity Fund3+18.6%
Counsel Canadian Growth Fund1+18.4%
Mackenzie Canadian Equity Fund1+17.9%
Sun Life BlackRock Cdn Equity Alpha Series A1+17.8%
IG Mackenzie North American Equity Fund1+17.4%
Desjardins Canadian Equity Fund1+17.1%
CI Canadian Investment Fund1+17.0%
Canada Life Cdn Focused Equity Fund Series F1+16.0%
Desjardins Canadian Equity Fund Series A1+16.0%
Leith Wheeler Canadian Equity Fund2+14.2%
Fidelity Canadian Opportunities Fund Series F2+14.1%
Canada Life Canadian Value Fund1+14.1%
Fidelity Canadian Opportunities Fund1+14.1%
Mawer Canadian Equity Fund Series A1+14.0%
Mawer Canadian Equity Fund1+14.0%
Dynamic Canadian Dividend Fund1+13.5%
MM Fund (unidentified — engine could not name the actual fund)1—
Rows marked "ETF returned as a fund" are exactly what they say: the prompt asked specifically for Canadian equity or Canadian focused equity mutual funds, filtered by CIFSC category, and the engine substituted an ETF instead — a different, unrequested product type presented as if it answered the question.
ETFTickerTimes Cited3-Year Return
BMO Equal Weight Global Gold Index ETFZGD5+53.7%
iShares S&P/TSX Global Gold Index ETFXGD7+53.7%
BetaPro Gold Bullion 2x Daily Bull ETFGLDU1+50.0%
BMO Junior Gold Index ETFZJG5+48.7%
Global X Global Semiconductor Index ETFCHPS1+47.5%
BetaPro NASDAQ-100 2x Daily Bull ETFQQU1+45.3%
Sprott Physical Silver TrustSVR.C3+41.9%
Horizons Gold Producer Equity Index ETFGLCC7+39.0%
Purpose Bitcoin ETF (fails Canadian-equity filter)BTCC.B2+38.9%
Fidelity Advantage Bitcoin ETF (fails Canadian-equity filter)FBTC3+36.8%
CI Galaxy Bitcoin ETF (fails Canadian-equity filter — most-cited overall)BTCX.B10+36.7%
Evolve Bitcoin ETF (fails Canadian-equity filter)EBIT2+35.6%
3iQ CoinShares Bitcoin ETF (fails Canadian-equity filter)BTCQ2+35.4%
BMO Covered Call Technology ETFZWT1+35.1%
BMO Equal Weight Global Base Metals ETFZMT2+33.3%
iShares S&P/TSX Capped Materials Index ETFXMA5+30.4%
Royal Canadian Mint CDN Gold Reserves ETFMNT3+30.3%
iShares Gold Bullion ETF (CAD-hedged)CGL.C2+30.2%
TD Global Technology Leaders Index ETFTEC4+29.9%
Evolve Canadian Banks Enhanced Yield ETFBANK1+28.6%
Invesco NASDAQ 100 Index ETF (CAD) (US index — fails filter)QQC2+26.1%
iShares S&P/TSX Capped Energy ETFXEG4+25.9%
iShares NASDAQ 100 Index ETF (CAD) (US index — fails filter)QQQ1+25.8%
TD Active U.S. Enhanced Dividend ETF (US equity — fails filter)TUED3+25.6%
iShares Core S&P 500 Index ETF (CAD-hedged) (US index — fails filter)XUS1+21.5%
iShares Core S&P/TSX Capped Composite ETFXIC1+21.5%
BMO S&P/TSX Capped Composite Index ETFZCN1+21.4%
Vanguard S&P 500 Index ETF (CAD) (US index — fails filter)VFV1+21.4%
BMO S&P 500 Index ETF (US index — fails filter)ZSP2+21.4%
Invesco S&P/TSX Composite ESG Index ETFESGC1+21.0%
Global X Crude Oil ETFHUC1+9.6%
CI Galaxy Ethereum ETF (fails Canadian-equity filter)ETHX.B2+5.5%
BetaPro Silver 2x Daily Bull ETFSLVD1−67.6%
Eleven of the 33 distinct ETFs cited across the study — crypto ETFs, US-listed index trackers, and leveraged products — do not meet the prompt's Canadian-listed, Canadian-equity filter at all. Several were cited more often than any of the five ETFs that actually met the filter and topped the verified return ranking.
2
Results by Engine, Trial by Trial
What each engine actually returned in each of its three independent sessions, per asset class

The tables below show each engine's top-five picks (or fewer, where an engine refused or ran out of answers) in each of its three trials, for each asset class. Colour marks how often a security recurred across that engine's own three trials — the same measure used throughout the Consistency findings above.

3 Same spot, all 3 trials
2 Appeared in 2 of 3 trials
1 Appeared in 1 trial
✕ ETF substituted, or a verified-wrong result
– No answer / refused
Equities
Trial 1Trial 2Trial 3
Celestica Inc.Celestica Inc.Celestica Inc.
Lundin Gold Inc.Bombardier Inc. (Class B)Lundin Gold Inc.
Cameco CorporationAritzia Inc.Cameco Corporation
Aritzia Inc.TFI International Inc.Shopify Inc.
Shopify Inc.Finning International Inc.Aritzia Inc.
Trial 1Trial 2Trial 3
Hammond Power SolutionsCelestica Inc.Hammond Power Solutions
Celestica Inc.Cameco CorporationSprott Inc.
Cameco CorporationAgnico Eagle Mines Ltd.Agnico Eagle Mines Ltd.
Lundin Gold Inc.Kinross Gold CorporationWheaton Precious Metals Corp.
Alamos Gold Inc.Wheaton Precious Metals Corp.Alamos Gold Inc.
Trial 1Trial 2Trial 3
Celestica Inc.Celestica Inc.HG Critical Capital
Bombardier Inc. (Class B)Bombardier Inc. (Class B)Hemisphere Energy Corp.
Aritzia Inc.Aritzia Inc.Fossil Mining Corp (unconfirmed)
Finning International Inc.Finning International Inc.Faraday Copper Corp
TFI International Inc.TFI International Inc.Titan Mining Corp
Trial 1Trial 2Trial 3
Celestica Inc.Energy Fuels Inc.Celestica Inc.
Bombardier Inc. (Class B)Celestica Inc.Cameco Corporation
Kraken Robotics Inc.Enerflex Ltd.IAMGOLD Corporation
Hammond Power SolutionsAritzia Inc.Suncor Energy Inc.
Aecon Group Inc.IAMGOLD CorporationImperial Oil Ltd.
Trial 1Trial 2Trial 3
Celestica Inc.Celestica Inc.Celestica Inc.
Bombardier Inc. (Class B)Bombardier Inc. (Class B)Bombardier Inc. (Class B)
Hammond Power SolutionsAritzia Inc.Aritzia Inc.
Lundin Gold Inc.TFI International Inc.No answer
TerraVest Industries Inc.MDA Ltd.No answer
Trial 1Trial 2Trial 3
No answer (refused)No answer (refused)Celestica Inc.
——Agnico Eagle Mines Ltd.
——Kinross Gold Corporation
——Lundin Gold Inc.
——Aritzia Inc.

Celestica's 14-of-18 citation rate is visible here at the individual-trial level: every engine except Perplexity named it, and ChatGPT and Grok named it in all three of their own trials. Grok's identical Trial 1 and Trial 2 picks, followed by a two-name Trial 3, and Copilot's byte-identical Trials 1–2 followed by an entirely unrelated Trial 3, are the same events described narratively in the Consistency section above — this is the underlying data they were drawn from.

Mutual Funds

Every mutual fund trial returned different names, engine to engine and often trial to trial. Some engines substituted ETFs outright; the source data also shows the Guardian Canadian Focused Equity Fund appearing under more than one series code within a single engine's own trials, reflecting ambiguity in how the engines resolved FundServ codes rather than a single consistent pick.

Trial 1Trial 2Trial 3
Guardian Canadian Focused EquityFidelity Special SituationsGuardian Canadian Focused Equity
BMO S&P/TSX Capped Composite (ETF)Guardian Canadian Focused Equity (series code variant)Guardian Canadian Focused Equity (series code variant)
Canada Life Cdn Equity Fund Series FCI Select Canadian Equity Series FFidelity Canadian Opportunities
Fidelity Canadian OpportunitiesRBC Canadian Equity Series FLeith Wheeler Canadian Equity Series F
IG Mackenzie North American Equity—Desjardins Canadian Equity Series A
Trial 1Trial 2Trial 3
Dynamic Power Canadian GrowthGuardian Canadian Focused EquityRefused
Mackenzie Canadian Equity FundCanada Life Canadian Value Fund—
Mawer Canadian Equity FundiShares S&P/TSX Capped Composite (ETF)—
Capital Group Canadian Focused EquityInvesco RAFI Cdn Fundamental (ETF)—
TD Canadian Small-Cap EquityVanguard FTSE Canada (ETF)—
Trial 1Trial 2Trial 3
RBC Canadian Equity FundRBC Canadian Equity FundCI Morningstar Momentum (ETF)
Capital Group Cdn Focused EquityCapital Group Cdn Focused EquityCI Morningstar Value (ETF)
Guardian Canadian Focused EquityGuardian Canadian Focused EquityDFA Canadian Vector Equity
Fidelity Canadian Growth FundDFA Canadian Equity FundRBC QUBE Canadian Equity
Dynamic Power Canadian GrowthCI Canadian Investment FundVanguard FTSE Canada All Cap (ETF)
Trial 1Trial 2Trial 3
Desjardins Canadian EquityScotia Canadian Growth Series AVanguard FTSE Canada (ETF)
Fidelity Canadian OpportunitiesRBC Canadian EquityiShares Core S&P/TSX Capped (ETF)
DFA Canadian Vector EquityMackenzie Canadian Equity FundBMO S&P/TSX Capped Composite (ETF)
DFA Canadian Core Equity Series ASun Life BlackRock Cdn Equity AlphaiShares S&P/TSX 60 Index (ETF)
NCM Core Canadian Fund Series ZMawer Canadian Equity FundMackenzie Canadian Equity Fund
Trial 1Trial 2Trial 3
Guardian Canadian Focused EquityGuardian Canadian Focused EquityGuardian Canadian Focused Equity
Scotia Canadian Growth Series AScotia Canadian Growth Series AScotia Canadian Growth Series A
Guardian Canadian Focused Equity (series code variant)—Guardian Canadian Focused Equity (series code variant)
Capital Group Cdn Focused EquityCapital Group Cdn Focused Equity—
—TD Canadian Small-Cap EquityTD Canadian Small-Cap Equity
Leith Wheeler Canadian EquityCounsel Canadian Growth FundRBC Canadian Equity
Trial 1Trial 2Trial 3
DFA Canadian Vector EquityGuardian Canadian Focused EquityCapital Group Cdn Focused Equity
RBC QUBE Canadian EquityDFA Canadian Vector Equity—
SEI Canadian Equity FundCapital Group Cdn Focused Equity—
Vanguard FTSE Canada All Cap (ETF)RBC QUBE Canadian Equity—
Vanguard FTSE Canada All Cap (ETF)MM Fund (unidentified)—
ETFs
Trial 1Trial 2Trial 3
MicroSectors Gold Miners 3XFidelity Advantage BitcoinBMO Equal Weight Global Gold Index
BMO Equal Weight Global Gold IndexPurpose BitcoinBMO Junior Gold Index
BMO Junior Gold Index3iQ BitcoiniShares S&P/TSX Global Gold Index
BetaPro Gold Bullion 2XEvolve BitcoinHorizons Gold Producer Equity Index
BetaPro Silver 2XCI Galaxy BitcoinEvolve European Banks Enhanced Yield
Trial 1Trial 2Trial 3
BMO Equal Weight Global Gold IndexBMO Equal Weight Global Gold IndexBMO Equal Weight Global Gold Index
BMO Junior Gold IndexBMO Junior Gold IndexBMO Junior Gold Index
iShares S&P/TSX Global Gold IndexiShares S&P/TSX Global Gold IndexiShares S&P/TSX Global Gold Index
Horizons Gold Producer Equity IndexHorizons Gold Producer Equity IndexHorizons Gold Producer Equity Index
iShares S&P/TSX Capped Materials IndexBMO Equal Weight Global Base MetalsiShares S&P/TSX Capped Materials Index
Trial 1Trial 2Trial 3
iShares S&P/TSX Capped Energy IndexiShares S&P/TSX Capped Energy IndexiShares S&P/TSX Capped Energy Index
CI Galaxy BitcoinCI Galaxy BitcoinCI Galaxy Bitcoin
Fidelity Advantage BitcoinVanguard S&P 500 IndexiShares NASDAQ 100 Index (CAD)
CI Galaxy EthereumBMO S&P 500 IndexTD Global Technology Leaders Index
Global X Crude OiliShares S&P 500 Index (CAD)BMO S&P 500 Index
Trial 1Trial 2Trial 3
CI Galaxy BitcoinCI Galaxy BitcoinCI Galaxy Bitcoin
Fidelity Advantage BitcoinGlobal X Global Semiconductor IndexBMO Covered Call Technology
3iQ BitcoinBetaPro NASDAQ 100 2x Daily BullInvesco S&P/TSX Composite ESG Index
Evolve BitcoinGlobal X NASDAQ 100 IndexiShares Core S&P/TSX Capped Composite
Purpose BitcoinTD Global Technology Leaders IndexBMO S&P/TSX Capped Composite
Trial 1Trial 2Trial 3
Sprott Physical Silver TrustSprott Physical Silver TrustSprott Physical Silver Trust
iShares S&P/TSX Global Gold IndexiShares S&P/TSX Global Gold IndexiShares S&P/TSX Global Gold Index
Horizons Gold Producer Equity IndexHorizons Gold Producer Equity IndexHorizons Gold Producer Equity Index
Royal Canadian Mint CDN Gold ReservesRoyal Canadian Mint CDN Gold ReservesRoyal Canadian Mint CDN Gold Reserves
iShares S&P/TSX Capped Materials IndexiShares S&P/TSX Capped Materials IndexiShares S&P/TSX Capped Materials Index

Grok's ETF picks are identical across all three trials — the exact pattern described in the Consistency findings as reflecting a memorised list rather than a live screen.

Trial 1Trial 2Trial 3
CI Galaxy BitcoinCI Galaxy BitcoinCI Galaxy Bitcoin
TD Active U.S. Enhanced DividendTD Active U.S. Enhanced DividendTD Active U.S. Enhanced Dividend
—TD Global Technology Leaders IndexTD Global Technology Leaders Index
—Invesco NASDAQ 100 Index (CAD)Invesco NASDAQ 100 Index (CAD)
—iShares Gold Bullion (CAD)iShares Gold Bullion (CAD)
CI Galaxy Ethereum——
iShares S&P/TSX Capped Energy Index (cited twice in Trial 1)——
This trial log is drawn from the underlying research paper's own trial-by-trial record. Where a fund or series code was ambiguous in an engine's original response — noted in several mutual fund trials above — that ambiguity is reproduced here rather than resolved, since resolving it would misrepresent what the engine actually returned.
3
Replication Parameters
Everything needed to run this test again, engine by engine, prompt by prompt

The parameters below are the complete set of conditions used to produce the 54 trials in this study. Holding all of them constant — same prompts, same order, same session discipline — is what makes a second run of this test comparable to the first one, even months later on different model versions.

  • Period: May 8–9, 2026 — Mississauga, Ontario, Canada.
  • Sessions: One fresh, private/incognito browser session per trial, opened immediately before the trial and closed immediately after the response was captured. No session was reused across trials, asset classes, or engines.
  • Trial count: Three independent trials per engine per asset class (18 trials per asset class, 54 trials total across all six engines).
  • Trial order: Within a given engine, one asset class was taken through all three of its trials before moving to the next — not interleaved. The exact sequence is worked through below.
  • Mode: Default consumer chat mode on each platform's standard web app. No deep research, extended thinking/reasoning, agent, or custom GPT/persona modes, and no developer API.
  • Account tier: Paid/pro consumer tier on every platform, to reflect how a financial advisor would actually use these tools.
  • Prompt wording: Identical wording, character for character, across every engine and every trial — see the five prompts below.
  • Verification baseline: Every returned security and figure checked against Buckler's verified market data as of April 30, 2026.

Every engine followed this same pattern: one asset class taken through all three of its trials, in a fresh incognito session each time, before moving to the next asset class.

StepAsset ClassTrialPrompt ChainSession
1EquitiesT1P1 → P4 → P5Fresh incognito
2EquitiesT2P1 → P4 → P5Fresh incognito
3EquitiesT3P1 → P4 → P5Fresh incognito
4ETFsT1P3 → P4 → P5Fresh incognito
5ETFsT2P3 → P4 → P5Fresh incognito
6ETFsT3P3 → P4 → P5Fresh incognito
7Mutual FundsT1P2 → P4 → P5Fresh incognito
8Mutual FundsT2P2 → P4 → P5Fresh incognito
9Mutual FundsT3P2 → P4 → P5Fresh incognito

At each step, the engine's full response to the core performance prompt was captured verbatim, followed immediately by its answers to the Prompt 4 sourcing audit and the Prompt 5 metadata request, all inside the same session. Nothing was summarized, cleaned up, or paraphrased at capture time. The same nine-step sequence was repeated for each of the remaining five engines.

EngineEquities TrialsMutual Fund TrialsETF TrialsTotal Trials
ChatGPT3339
Claude3339
Copilot3339
Gemini3339
Grok3339
Perplexity3339
All Six Engines18181854
AI Engine Retrieval Engine Version Knowledge Cutoff Notes for Replication
ChatGPTBing-based browsing toolGPT-5 or GPT-4oEarly-to-mid 2024Standard ChatGPT web app, paid tier, browsing on. Default chat interface, not GPTs/agents.
ClaudeAnthropic's web_search + web_fetch toolsClaude Sonnet 4.xLate 2024 / May 2025Standard Claude.ai web app.
CopilotBing-based browsingMicrosoft Prometheus stack (GPT-4o / GPT-4 Turbo)Late 2023 to mid 2024Standard Copilot Chat (consumer), Creative/Balanced mode default.
GeminiGoogle Search toolGemini 2.5 ProLate 2023 to mid 2024Gemini consumer app. Audited as not using the Google Search tool or live external databases across all nine trials.
GrokxAI web searchGrok 4Late 2023SuperGrok, paid subscription.
PerplexityPerplexity's own crawler + web indexSonar (or Sonar Large)Late 2023 to mid 2024Default Perplexity chat mode, not Pro Search/Copilot Search.

Each session ran one asset-class chain, opening prompt first: Equities = Prompt 1 → Prompt 4 → Prompt 5. Mutual Funds = Prompt 2 → Prompt 4 → Prompt 5. ETFs = Prompt 3 → Prompt 4 → Prompt 5. No follow-ups, rewording, or clarification requests were used at any point.

"As of your most recent data, identify the top five Canadian equities meeting the following criteria. State explicitly which as-of date(s) you used."Prompt 1 — Equities

Filters: issuer domiciled in Canada; market capitalization greater than $1 billion CAD, ranked by 12-month trailing total return. Required six data points per security (market cap, share price, trailing P/E, trailing 12-month dividend yield, 3-year price appreciation, 3-year beta vs. S&P/TSX Composite), single table, ticker/issuer/exchange included, as-of date stated, no unsolicited commentary.

"As of your most recent data, identify the top five Canadian equity or Canadian focused equity mutual funds. State explicitly which as-of date(s) you used."Prompt 2 — Mutual Funds

Filter: CIFSC category of Canadian Equity or Canadian Focused Equity, ranked by 3-year annualized return. Required nine data points per fund (fund name, FundServ code, MER, 3-year and 5-year annualized return, 5-year standard deviation, maximum drawdown, AUM, CIFSC category), N/A for unavailable data rather than declining to answer.

"Identify the top five Canadian-listed ETFs ranked by 3-year annualized return (highest to lowest)."Prompt 3 — ETFs

Filter: listed on a Canadian exchange. Required seven data points per ETF (ticker, issuer, MER, AUM, 3-year annualized return, 1-year return, average daily volume), same table and N/A conventions as Prompt 2.

"Please enumerate all the sources you cited or relied on in your previous response. For each source, provide: (1) the publisher or site name, (2) the specific data point or claim it supported, (3) the full URL, (4) the as-of date of the figure."Prompt 4 — Enumeration and Provenance Disclosure

Run immediately after the core prompt in every session. Required the engine to classify each source as primary (issuer fact sheets, KYP documents, Bloomberg, FactSet, Morningstar Direct, S&P Capital IQ, exchange filings, SEDAR+) or context-only (guidance articles, blogs, Wikipedia); to state whether each URL was actually fetched, used only as a search-snippet reference, a placeholder, or never retrieved at all; to give a confidence rating per data point (HIGH/MEDIUM/LOW); to confirm directly whether Bloomberg, FactSet, Morningstar Direct, S&P Capital IQ, issuer KYP/fact-sheet PDFs, SEDAR+, and exchange-level feeds were accessed; and to disclose retrieval failures. Closed with: "If you cannot answer any portion of this disclosure honestly, say so directly rather than inferring."

"For research methodology and audit purposes, please document the technical metadata of this session."Prompt 5 — Session Metadata

Ten items requested per session: date/time and timezone as reported; model name and exact version; knowledge cutoff; whether live browsing/search was enabled and which retrieval engine; whether any deep-research, reasoning, or extended-context mode was active; custom instructions, system prompt, persona, or memory settings in effect; inferred or stated geographic/regional session settings; a session/conversation/response ID if available; tool restrictions, rate limits, or capability constraints; any other metadata a replicator would need. Closed with: "If you cannot answer any of the above, state so explicitly rather than guessing."

  • Consistency: an engine's three trials for a given asset class compared against one another — same securities, same order, same figures, or not.
  • Accuracy: securities and return figures compared against Buckler's verified market data as of the trial period; citation rate calculated as the count of trials (of 18 per asset class) in which a security was named, expressed as a percentage.
  • Transparency: the engine's own Prompt 4 and Prompt 5 answers checked against what it had actually done — whether a claimed source was real and fetched, a snippet reference, a placeholder, or fabricated outright.
Expect engine versions, retrieval tooling, and knowledge cutoffs to have moved on since May 2026. Documenting them this precisely is what lets a replication attempt record what changed rather than requiring every variable to hold identical — the prompts, session discipline, and scoring method above are what actually needs to stay fixed.
This is the kind of review Buckler documents automatically — every security, every trigger, every time. See how →