Comparing ChatGPT, Bard, and Amazon Q: A Feedback Analysis
A practical comparison of ChatGPT, Bard, and Amazon Q across general, technical, and entrepreneurial prompts, including strengths and trade-offs.

Lightweight model evaluation frame
categories:
- general_life_guidance
- technical_developer_use_cases
- entrepreneurial_strategy
scoring_dimensions:
- completeness
- specificity
- actionability
- consistency
questions_per_category: 4
prompt_engineering: disabled
A simple framework for comparing assistant behavior with minimal evaluator bias.
Original publication
This post is adapted from the original DEV Community article: Comparing ChatGPT, Bard, and Amazon Q: A Feedback Analysis
Why this comparison is useful
Choosing an assistant is less about brand preference and more about fit-for-purpose performance. This comparison tested three assistants on twelve prompts across three user segments without prompt engineering to observe baseline behavior.
Evaluation setup
The prompts were grouped into:
- Non-technical general-purpose questions.
- Technical and developer-focused questions.
- Entrepreneurial and go-to-market questions.
Each category used four prompts to balance breadth and consistency.
Observed patterns
Bard (now Gemini)
- Tended to provide broad and comprehensive responses.
- Often included additional context helpful for exploration.
Amazon Q
- Performed strongly on technical and developer-centric prompts.
- Needed broader depth for non-technical and entrepreneurial contexts in this sample.
ChatGPT
- Returned answers across all categories.
- Response quality benefited from tighter follow-up for concision and precision.
Practical implications for teams
- If your workload is heavily technical, domain-tuned assistants may feel stronger by default.
- If you need broad ideation and cross-domain responses, generalist models can be more flexible.
- For business-critical usage, create your own prompt set and evaluate assistants against your exact tasks.
Recommended internal benchmark flow
- Define categories from real team workflows.
- Use fixed prompts and fixed scoring dimensions.
- Evaluate quality with at least two reviewers.
- Re-run quarterly because model behavior evolves quickly.
Caveats
This analysis is directional, not absolute. Results depend on prompt wording, model version, and update cadence. Any assistant decision should be validated in your own operating context.
Read more
For full prompt lists and the original comparison details, read the source article: