How we know

Research methodology

RoomBrew makes factual claims about thousands of hotels. This page explains exactly where those claims come from, how much weight each kind of evidence carries, and what we do when the evidence runs out.

The rule everything else follows from

We do not publish a fact we cannot trace to a source. Every coffee-related claim on a hotel page is stored against the page it came from, the sentence that said it, and the date we read it. If we cannot do that, the field stays unknown.

This has an obvious cost: large parts of the database say “not yet confirmed”. We think that is the right trade. A confident wrong answer about a coffee machine sends someone to buy capsules that do not fit, and quietly destroys the reason to trust anything else on the page.

Where hotels come from

Hotel identity — name, address, coordinates, brand, phone number and official website — comes from OpenStreetMap, used under the Open Database Licence and attributed on every hotel page. Editors can also add a property by hand.

A hotel exists in RoomBrew because it exists in one of those sources. We never generate properties to inflate coverage.

Where coffee facts come from

We rank sources into three tiers, and the tier determines how much weight a claim carries.

Tier 1 — primary

The hotel’s own website and its chain’s official pages: room descriptions, amenity lists, FAQs, press releases about renovations. This is the strongest evidence we can get without standing in the room, because the hotel is describing its own product.

Tier 2 — strong secondary

Recent room-tour video, travel journalism, reputable travel writing, dated guest photography. Useful, and often more current than an official page that has not been updated since a refurbishment.

Tier 3 — discovery leads

Forum posts, community discussion, question-and-answer pages. These are good at telling uswhere to look and poor at settling anything on their own. A Tier 3 claim never reaches a confident status by itself.

How we read a page

Our extractor is deliberately literal. It looks for specific phrases and records the sentence containing them, and it declines several inferences that would be easy and wrong:

  • A coffee table is not a coffee maker.
  • A Keurig described in the lobby section is lobby coffee, not in-room coffee. Claims about the room have to come from text about the room.
  • A machine is never inferred from the hotel’s brand, class or price. Franchised properties within one brand differ enormously.
  • Nespresso is never resolved to Original or Vertuo unless a source says which. The two are physically incompatible, and guessing would send travellers to buy capsules that will not fit.
  • A page about one class of room describes that class. A finding from a suites page is recorded against the suites, not the whole property.

Where a page is ambiguous, the extractor stays silent. Silence is a correct answer; a guess is not.

How sources are accessed

We fetch pages under a set of rules that are not negotiable internally:

  • We read robots.txt before the first request to any host and obey it, including Content-Signal directives where a site publishes them. A site that signals it does not want to be indexed is not indexed.
  • We identify ourselves honestly in every request, as RoomBrewBot, with a link back to this page.
  • We space requests to a single host and read only a handful of pages per hotel. We are a small reference index and have no business adding measurable load to anyone’s site.
  • If a site refuses us, we stop. An HTTP 401, 403 or 429 is treated as a decision, not an obstacle. We do not disguise the crawler, rotate addresses, or work around technical access controls, and several major hotel chains decline automated access as a result. Where that happens, the hotel page says so plainly rather than implying we found nothing.
  • We store a short excerpt as proof of a claim — enough to verify it, far short of republishing anyone’s page.

Resolving disagreement

Sources contradict each other constantly, usually because a hotel changed machines or because two pages describe different room types. When that happens we do not average the answers or quietly pick one.

A newer, equally authoritative source supersedes an older one — hotels really do swap suppliers, and the recent page is usually the current truth. Where two sources are comparably strong and comparably recent, the fact is marked conflicting, both readings stay visible, and the page tells you the reports differ.

RoomBrew Confidence

The score on each hotel page is a practical trust indicator, not a statistical claim. It is built from things we actually recorded:

  • Source tier — an official page starts higher than a forum post.
  • Independent corroboration — two pages on the same domain count as one source.
  • Freshness — evidence decays because hotels change. A three-year-old room tour is still evidence, just weaker.
  • Traveller confirmations — someone who was in the room last month is excellent evidence.
  • Conflicts — unresolved disagreement lowers the score substantially.
  • Editor review — a human who has checked the claim against the source outranks the model.

Every hotel page shows the reasons behind its score, not just the number. If the reasons do not persuade you, the number should not either.

Currentness

Two dates matter and they are different. Last checked is when we last went and looked. Evidence dated is how old the underlying source is. A page checked yesterday against a 2021 amenity list is not current information, and we say so.

Evidence older than about eighteen months marks a fact stale, and the page warns you the setup may have changed. Busy properties are re-checked more often than long-tail ones.

History

When a fact changes, the old value is kept. A hotel that replaced its Keurigs with Nespresso machines shows both eras on its page, with dates. We never silently overwrite the past — partly because it is useful, and partly because it is the only way to audit ourselves.

Traveller reports

Anyone can confirm or correct a hotel’s setup, without an account. Reports are reviewed by a moderator before they affect anything displayed, and a report about a stay last month carries more weight than one about a stay three years ago.

We store no personal information with a submission. The submitter’s address is kept only as a salted hash, used for rate limiting.

What we are not good at

An honest list, because every methodology page should have one.

  • Coverage is uneven. We seed the busiest travel markets first. A hotel in a smaller city may sit unresearched for a long time.
  • Some chains are closed to us. Several large hotel groups decline automated access, and we respect that. For those properties, traveller reports are the main route to a confirmed answer.
  • Room-type detail is thin. We capture it where a source is explicitly about one class of room, which is a minority of the time.
  • Official pages go stale. A hotel that refurbished last year may still be described by a page written before it. This is what the freshness weighting and the traveller reports exist to catch.

Where things currently stand

RoomBrew currently tracks 17,896 hotels across 27 countries and 57 cities. 3,628 have had at least one research pass, drawing on 7,309 sources. 405 have a confirmed in-room coffee maker and 2 are confirmed to have none.

These numbers are generated from the live database, not written by hand, and they will go up and down as research runs.

Corrections

If something here is wrong, we want to know and we will fix it. See our corrections policy, or use the confirmation form on any hotel page.