[verdict][curated prompt][source: 100questionsai.com]
Can I vibecode 100 Questions?
// buildable in a weekend, but real gaps stay open
A personal CLI that asks the same questions across four model APIs and compares the answers is weekend-buildable, but matching the product's web-grounded runs, source normalization, failure handling, durable evidence, scoring, and polished reports takes substantially more work.
confidence: high
what it is: Evidence-linked brand visibility benchmarks across OpenAI, Claude, Gemini, and Grok
Buildability index · an editorial game
- Price 9 $ · one-time weight: plus 1
- Time multi-day weight: plus 1
- Category seo and marketing no weight
- Moat execution quality · integrations weight: minus 2
- Confidence high weight: plus 1
- What you lose 5 items weight: minus 2
- Site 100questionsai.com (si apre in una nuova scheda) no weight
They cancel out. This is where it comes down to how much the subscription annoys you.
Gioco editoriale: il verdetto dice se un agente può, l’indice se conviene.
What you build
Generate buyer questions, run each through four web-grounded model APIs, detect brand and competitor mentions, collect citations, and render a comparison report.
what you need
- OpenAI, Anthropic, Google Gemini, and xAI API keys
- provider-specific web search or grounding tools
- durable run storage
- URL and citation normalization
- report generation
The personal core is approachable; production-grade evidence and cross-provider consistency are the hard parts.
The prompt
A weekend with a coding agent. The gaps that stay are right below, under “what you lose”.
Build me a local AI visibility benchmark for one brand. Requirements:
- Use Node 22, TypeScript, official provider SDKs, SQLite, and a CLI.
- `benchmark --domain example.com --description "..."` creates one immutable run.
- Generate 25 buyer questions from the domain and description, or accept a JSON question file.
- Ask the exact same questions through OpenAI, Anthropic, Gemini, and xAI.
- Use each provider's supported web-search or grounding tool; keys live only in `.env`.
- Limit concurrency per provider, retry transient failures, and preserve failed cells in the report.
- Store prompts, raw answers, citations, timestamps, model ids, and errors in SQLite.
- Detect exact and case-insensitive brand mentions; allow aliases in a config file.
- Extract named competitors with one structured LLM pass after all answers are stored.
- Normalize citation URLs by hostname, canonical URL, and stripped tracking parameters.
- Compute visibility by provider, answer coverage, owned-domain citation rate, and top sources.
- Show missed questions where competitors appear but the target brand does not.
- Render a self-contained static HTML report with filters and expandable raw evidence.
- Export questions, answer metrics, competitors, and citations as CSV files.
- Every aggregate metric must link back to the answer rows used to calculate it.
- Out of scope: accounts, billing, teams, scheduled monitoring, and recommendation generation.
- Include fixture-based tests for mention detection, URL normalization, and metric calculations.
- README: setup, provider-specific grounding caveats, estimated API cost, and exact run commands. Curated prompt: written and reviewed by hand for this app. In English on purpose — it is the language coding agents work best in.
What you lose
- reliable orchestration and retries across four providers
- normalized citations and evidence-linked metrics
- competitor and missed-question extraction
- stored point-in-time reports and comparisons
- polished exports and action recommendations
Why people still pay
They pay for a repeatable, frozen benchmark with provider failures handled, citations normalized, every metric tied to evidence, and a report that is ready to act on.
moat: Execution quality Integrations what a moat is
workflow/reliability/evidence normalization
Free alternatives
Not in the mood to build it? These already exist, they are free or open source, and we checked them one by one.
Rejected (3) — and why
- GEO/AEO Tracker (si apre in una nuova scheda) — A polished dashboard with 221 stars and 28 commits; every real run still depends on Bright Data's commercial scraper.
- OneGlanse (si apre in una nuova scheda) — The feature list is dead-on; 141 stars, no releases, five provider logins, three databases and a residential proxy are not a non-builder product.
- OpenSEO (si apre in una nuova scheda) — Mature and installable, but its visibility layer stops at ChatGPT and Google AI Overview; Claude, Gemini and Grok are the product here.
Do you agree?
The vote balance
Ancora nessun voto: il tuo è il primo.
Nessun voto ancora — il primo pesa.