Comparing ChatGPT, Bard, and Amazon Q: A Feedback Analysis

A practical comparison of ChatGPT, Bard, and Amazon Q across general, technical, and entrepreneurial prompts, including strengths and trade-offs.

Updated Feb 17, 2024#generativeai#chatgpt#bard#amazonq#evaluation
Cover visual for chatbot comparison across ChatGPT, Bard, and Amazon Q

Original publication

This post is adapted from the original DEV Community article: Comparing ChatGPT, Bard, and Amazon Q: A Feedback Analysis

Why this comparison is useful

Choosing an assistant is less about brand preference and more about fit-for-purpose performance. This comparison tested three assistants on twelve prompts across three user segments without prompt engineering to observe baseline behavior.

Evaluation setup

The prompts were grouped into:

  1. Non-technical general-purpose questions.
  2. Technical and developer-focused questions.
  3. Entrepreneurial and go-to-market questions.

Each category used four prompts to balance breadth and consistency.

Observed patterns

Bard (now Gemini)

  • Tended to provide broad and comprehensive responses.
  • Often included additional context helpful for exploration.

Amazon Q

  • Performed strongly on technical and developer-centric prompts.
  • Needed broader depth for non-technical and entrepreneurial contexts in this sample.

ChatGPT

  • Returned answers across all categories.
  • Response quality benefited from tighter follow-up for concision and precision.

Practical implications for teams

  • If your workload is heavily technical, domain-tuned assistants may feel stronger by default.
  • If you need broad ideation and cross-domain responses, generalist models can be more flexible.
  • For business-critical usage, create your own prompt set and evaluate assistants against your exact tasks.
  1. Define categories from real team workflows.
  2. Use fixed prompts and fixed scoring dimensions.
  3. Evaluate quality with at least two reviewers.
  4. Re-run quarterly because model behavior evolves quickly.

Caveats

This analysis is directional, not absolute. Results depend on prompt wording, model version, and update cadence. Any assistant decision should be validated in your own operating context.

Read more

For full prompt lists and the original comparison details, read the source article:

Read the full feedback analysis on DEV