Methodology

How this is measured

This page exists to be checked. If the method is unclear or looks gameable, that's a problem with the method, not with your reading of it. Email me and I'll fix it.

The five engines

Every question is asked against all five of the following, unmodified:

  • ChatGPT
  • Claude
  • Gemini
  • Perplexity
  • Google AI Overviews

No custom instructions, memory, or account personalization beyond what's unavoidable for a logged-in session. Four of the five, ChatGPT, Gemini, Perplexity, and Google AI Overviews, are reached fully logged out: no account, no memory, nothing to disable. Claude has no anonymous or guest mode at all, so it's the one exception: a real account is used, with memory and personalization turned off and existing memories cleared before each run. That asymmetry isn't a methodology choice, it's what each platform actually offers a real buyer, and it's disclosed here rather than smoothed over.

Runs are conducted from a US IP address, unset by any VPN or proxy. Location can affect trades-software answers, so it's pinned and disclosed rather than left to chance: the August 2026 run was conducted from Atlanta, GA. Any change in location for the October or December runs will be disclosed here.

Repetitions and sessions

Each of the 21 questions is asked 5 times per engine, and each of those 5 repetitions starts in a fresh session: new chat, no prior context in that session. This is because a single answer from a single session says nothing about whether an engine reliably names a company or whether that run was noise. Presence is reported as a rate across the 5 repetitions (e.g. "named in 4 of 5 runs"), not as a single yes/no.

That's 21 questions times 5 engines times 5 repetitions: 525 logged responses per run, before the brand-direct loop below.

This question set deliberately over-samples solo and small-shop buyers. That's where the tracked challengers position themselves, and it's the segment where an AI recommendation can plausibly drive an actual purchase: smaller operators buy self-serve off published pricing, while larger contractors buy through demos, quotes, and procurement, where an AI answer influences a shortlist at most. The skew is a choice, not an oversight.

The three measures

  • Presence: which companies are named in the response, and how often across the 5 repetitions. A company that's never named for a given question is recorded as absent, not skipped.
  • Accuracy: for each company named, whether what the engine says about it (pricing, feature claims, positioning, who it's for) is verifiably correct against that company's own current public materials. Errors are logged with what was said and what's actually true.
  • Sourcing: which URLs the engine cites or links to (where the engine exposes citations at all), and, where an inaccurate claim can be traced to a specific source, which page it came from.

The 21 questions

Grouped by where the buyer is in their decision, bucket A through D. The full set is frozen for the life of the project. See below for why.

Bucket A: Problem-aware: the buyer hasn’t named the category yet

Does the category get recommended at all, or does the model suggest a spreadsheet?

  1. How do I stop losing track of jobs and invoices in my electrical business
  2. How can a small electrical contractor schedule jobs and dispatch techs
  3. How do electricians write estimates faster
  4. What's the best way to handle missed customer calls as an electrician

Bucket B: Category-aware: actively evaluating

The core presence test. Who makes the shortlist.

  1. Best software for electrical contractors
  2. Best field service software for a small electrical shop
  3. What software do electricians use to run their business
  4. Best electrical contractor software for a 5-person shop
  5. Best software for a solo electrician
  6. Best electrical estimating software

Bucket C: Capability-specific: deserve-to-be-named tests

Each describes a tracked company's actual product. If an incumbent gets named instead, that's the finding.

  1. AI estimating software for electricians
  2. Software built specifically for electrical contractors
  3. AI phone answering service for electrical contractors
  4. Electrical contractor software with load calculations and panel schedules
  5. Software for electrical contractors that handles permits and code compliance

Bucket D: Comparison and decision: where a challenger has its best structural shot

  1. ServiceTitan alternatives
  2. Cheaper alternative to ServiceTitan for electricians
  3. Jobber vs Housecall Pro for electrical work
  4. Is ServiceTitan worth it for a small electrical contractor
  5. How much does electrical contractor software cost
  6. I'm outgrowing Jobber, what should an electrical contractor upgrade to

    The upgrade moment is where a challenger has its single best structural shot at being named: the buyer has already rejected the entry-level option and is explicitly asking to be told something new. If the models answer with ServiceTitan every time, that's a finding in its own right: the ladder has exactly two rungs.

Brand-direct loop

Separately from the 21 questions above, every tracked company is asked about directly, by name, using the two questions below. Each runs 3 repetitions per engine, same fresh-session rule, not counted toward the 21 total:

  • What is [company]
  • Is [company] good for electrical contractors

This isolates accuracy from presence: a company can be described correctly when asked about directly while still never surfacing in the category questions above, and that gap is itself a finding.

Run schedule

RunDate
1August 12, 2026
2October 2026
3December 2026

Three runs across five months against an identical question set turn single snapshots into a trend: whether a company's presence is stable, growing, or disappearing, and whether accuracy improves as engines update.

Why the question set is frozen

If the questions changed between runs, any difference in results could come from the new questions instead of from the engines changing. Freezing the set is what makes "run 2 vs. run 1" a comparison of the same thing measured twice, rather than two different measurements. The tradeoff is real: the set can't be expanded or corrected mid-project without breaking that comparability, so the 21 questions and the company list above are final for August, October, and December 2026, frozen as of August 5, 2026. Anything learned about gaps in the question set gets written up and applied to a future, separately versioned project, not folded back into this one.

What this doesn't claim

This isn't a ranking, a benchmark of model quality, or an endorsement of any company. A company being named more often isn't necessarily "better," it may just be better indexed, more heavily reviewed, or older. The companies page lists what's tracked without characterizing any of them, and it stays that way until there's data to back a claim up.

The full raw responses, coded by hand against the definitions above, publish as CSV and JSON after each run. See the data page.