Practical AI for Australian Small Business

Methodology: How We Research, Source and Score

Last updated: 8 September 2026

Why this page exists

Most pages on this site carry a score, a verdict, or a statement about what a regulator has said. This page explains how those are produced, what evidence sits behind them, and where the limits are, so you can decide for yourself how much weight to give them.

It also states plainly what this site does not do. That part matters as much as the rest.

What we are, and what we are not

We are a research and comparison site for Australian small and medium businesses working out where AI tools fit. We read vendor documentation, pricing pages, regulator guidance and public evidence, and we write up what we find.

We are not a law firm, an accounting practice, or a compliance auditor. On any governance or compliance topic we act as a navigator rather than an interpreter. That means we explain that an issue exists, point to where the authoritative guidance lives, summarise what that guidance says with attribution and a link, and direct you to the relevant regulator or your own adviser for anything specific to your situation.

Concretely, we do not tell you what your legal obligations are. We report what a named authority has published, and we link the document so you can read it yourself. If you see a sentence on this site that states an obligation as fact rather than attributing it to a source, that is a mistake and we would like to hear about it.

How we research

Our work is desk research and evidence triangulation. We read primary documents: the regulator's own guidance, the vendor's own pricing and documentation, and published terms. Where a claim can be checked against a named document, we check it against that document rather than against another article about it.

We do not run multi-week deployments inside live businesses, and we do not present our findings as though we had. Where a review reflects use of a product, we say so. Where it reflects documentation and public evidence, we say that instead.

Every consequential factual claim is tied to the specific source document it rests on, and that link is recorded rather than left implicit. On our reference pages the standard is section level: if a section attributes something to an authority, the source link appears in that section, not merely somewhere else on the page.

What "verified" means here, and what it does not

When we say a claim is verified, we mean a person or a checking process has read the cited document and confirmed that it says what we report it as saying. It is a check of faithfulness to the source.

It is not a check that the source is correct. If a regulator's guidance is ambiguous, or a vendor's pricing page is wrong or out of date, a faithfully verified claim can still mislead. We are reporting the record, not auditing it.

Verification is also not complete. A minority of the claims in our register are recorded as unverified or as resting on a document we could not retrieve, and they are marked that way in our own records rather than quietly counted as done.

Reference pages

Fourteen pages on this site are reference pages. They were written to be cited by other people, and they are held to a stricter standard than the rest of the site. Each one carries a note at the foot recording what was checked on it, and when.

The standard has four parts:

  • Every consequential claim links the primary document. Not a summary of it, and not another article about it, but the regulator's or the legislature's own text wherever one exists.
  • Each source was fetched and read. The claim was matched against the words in the document. The research of the model that drafted a page was never accepted as evidence for that page's own claims.
  • A claim that could not be confirmed is recorded as unverified. It is not dropped, and it is not quietly counted as done. Six claims across the fourteen pages stand in that state because the source host could not be reached when we checked.
  • The check is written down. Every one of the fourteen has a dated verification record naming the claims checked and the result. Those records are why the figures in each page's note can be re-derived rather than simply asserted, and an automated check re-derives them.

Across the fourteen pages, 132 claims were checked. 123 were confirmed against the cited source, 6 are recorded as unverified, and 3 describe how this site sourced its own comparison and were checked against that working instead of an outside document. Seven errors were found and corrected before the pages were published.

What this standard does not do. It is desk research against published documents. Nobody here deployed the products described or took legal advice on the regulations. Verification checks that we reported a document faithfully. It does not check that the document is correct, and a document can be amended or withdrawn after we read it. It also applies to these fourteen pages only, not to the rest of the site. Where we later find an error, it is listed in our corrections log.

Known limits

These are the weaknesses we know about. They are listed because a methodology page that only describes strengths is marketing.

  • Some cited documents have gone offline. Regulators and vendors move and retire pages. A number of sources in our register no longer resolve, and some of those carry claims that appear in live articles. We track which ones, and a dead source is a queue item rather than something we treat as still checked.
  • Some sources cannot be retrieved for checking at all. Certain documents sit behind blocks or return no readable content to an automated fetch. We treat that as a fact about our own access rather than about the document, so it is recorded as unverifiable, not as verified and not as dead, and it is re-tested later through a different route before we draw any conclusion from it. That matters more than it sounds: in September 2026 a re-check of every source we had recorded as unreachable found that most of them were readable after all, and six claims we had parked as unverifiable were confirmed word for word against the document we had failed to open.
  • Pricing goes out of date faster than we can re-check it. Vendors change prices without notice, and many pricing pages now return no figures to any automated check. Always confirm current pricing with the vendor before making a decision.
  • Our coverage is uneven. We have written far more about some categories than others. A tool being absent from this site is not a judgement about it.
  • We are a small operation. There is no testing lab and no panel of reviewers behind these pages.

The NTK Score

Tools we score are rated across five pillars, each worth up to 20 points, for a total out of 100:

  • Relevance: how well the tool fits the problem a small business is actually trying to solve.
  • Effort: how much setup, training and ongoing admin it takes to get value from it.
  • Adoption: how likely staff are to use it day to day, not just trial it once.
  • Commercial value: what it costs against what it saves or earns, at small-business scale.
  • Trust: vendor stability, data handling transparency, pricing stability, support reputation and track record.

Each pillar carries a written one-sentence rationale, a named evidence source, an evidence tier recording whether that source is a primary document, the vendor's own published statement, or a weaker public signal, and a confidence level. A pillar score with no rationale or no named source is rejected before the page can be saved.

The verdict tiers are: Recommended at 80 or above, Recommended with caveats at 65 to 79, Pilot first at 50 to 64, and Not recommended below 50.

A weak Trust pillar caps the verdict regardless of the total. For a general small-business assessment, a Trust score below 12 out of 20 caps the verdict at "Pilot first". For tools aimed at regulated industries such as healthcare, legal or financial services, the bar is higher and the cap applies below 14 out of 20, because the cost of a trust failure in those settings is greater. The cap can only lower a verdict, never raise one.

Separately, a confirmed material concern can force a "Not recommended" verdict regardless of the numeric score. This is used rarely and requires a written reason.

Scores also carry caveats: specific, named concerns that may affect some readers and not others. A caveat can change who we would recommend a tool to without changing the number, and caveats never move the score on their own.

Every score records the date it was evaluated and a date for its next review. A score is a snapshot of a moving product, and the date is there so you can judge how stale it might be.

Scores you may have seen before

Earlier reviews carried an interim star rating out of 10 alongside the NTK Score. It was a same-session editorial judgement rather than an evidence-backed score, it was meant to be replaced once a product went through the full process, and on some pages it was not. In September 2026 we removed every remaining interim star rating. Those pages now show one score, the NTK Score, and nothing else. Where the two had disagreed, the NTK Score is the one we stand behind.

Corrections

We correct errors rather than quietly editing them away. A correction is an addition to the record, not a replacement of it. Our corrections policy explains what we treat as a correction and how to report one.

When a factual error is found in one place, we check whether the same claim appears elsewhere on the site and correct every instance, rather than only the page that was reported. A wrong figure is rarely wrong in only one article.

Independence

We are not affiliated with any AI vendor, software company or consultancy, and we do not accept payment in exchange for reviews or editorial coverage. No vendor sees a review before it is published, and no vendor can pay to change a score or a verdict.

Where affiliate links appear they are disclosed clearly, and a commercial relationship is never a reason for a recommendation. Our independence and disclosure policy sets out the full position.

How this stays current

Automated checks run against the live site on a schedule. They re-check that cited sources still resolve, flag claims whose supporting figures may have moved, verify that pages carry the structure and attribution they are supposed to, and check this site's own public statements against what the site actually does.

Those checks produce a queue for a person to work through. They do not silently rewrite published content, and a check that cannot reach a source records that it could not, rather than recording a pass.

Telling us we got something wrong

If something here is inaccurate, out of date or misattributed, please tell us through the contact page with a link to the page and a description of the problem. Corrections that come from readers are welcome and are handled the same way as ones we find ourselves.