Every major search engine today (Google, Bing, Yahoo, Baidu, Yandex) is free to use and funded almost entirely by advertising. The ad model has been, and currently is, the most successful way a search engine is able to be profitable. However, in it’s early history this wasn’t always the case. Subscription models were tried and failed, and the ad model wasnt perfect at first and had to be refined. In the end the now common search engine ad model is the result of historical consumer, legal, technical, and economic events that played out between 1990 and early 2000’s.
This article walks through the search engines emergence, from it’s origins of failed attempts to organize the web, and a demand problem that forced into existence much of what we see today. We cover the methods attempted to commercializing it, including ones that failed, and why the ad model ended up being the inevitable victor.
Moreover, this article serves as a pretext to understanding search engines and why they are the way they are, and what value-systems the businesses running them have. Value-systems that ultimately determine why they rank what they rank, and what the algorithms they use are actually trying to achieve. In short, a search engine isn't entirely a neutral tool, though it optimizes for it in certain ways. What it shows, what it hides, and what it rewards all trace back to how it's funded and who it has to satisfy.
The Internet That Banned Advertising (NSFNET)
In order to fundamentally understand search engines, it's worth going over some early internet history, where they emerged. Originally the National Science Foundation Network (NSFNET), was the government backbone that funded much of the early internet, including it’s research and development. In it’s earliest form, the internet was an infrastructure web that connected American universities and research centers through the late 1980s and early 1990s, operating under an “Acceptable Use Policy” (AUP). NSF funded the network strictly for scientific research and education, and the AUP policy reflected that directly.
The Acceptable Use Policy: What NSFNET Actually Prohibited
NSFNET existed to connect researchers and educators to the supercomputing centers it funded, and to each other. Its Acceptable Use Policy spelled out, in plain terms, what fell outside that purpose. Advertising was named explicitly, and so was fundraising and personal for-profit ambitions. Even extensive personal or private use, anything beyond occasional, non-business use, was listed as unacceptable.
A Playground for Technologists: Emtage on the Pre-Commercial Internet
One person who lived through this period, Alan Emtage, then a systems administrator at McGill University, described the culture of the time directly: people were exploring the technology cooperatively, with "not much thought to making money off of it." He's also noted, however, that it was already clear to people that commercial activity was coming, eventually.
Archie and the First Internet Search
In 1989, Alan Emtage was a systems administrator at McGill University, tasked with manually searching public FTP servers for software useful to students and faculty. After realizing how impractical doing this manually was, he automated the process instead.
The result, built with colleagues Peter Deutsch and Bill Heelan, was Archie, named after "archive" with the v dropped. Archie predates the world wide web, and it also indexed FTP servers rather than web pages. Site administrators registered their FTP archive with the service, and a script then logged in on a schedule, pulled the directory listing, and folded it into a searchable database of file names, without looking at the contents of the files themselves.
What Archie automated was the repetitive part of the job, which was requerying known servers on a schedule instead of by hand. What it didn't automate was the discovery itself. Nothing went looking for FTP servers that hadn't been registered. Anyone who wanted Archie to index their FTP files, had to submit them. Emtage has referred to Archie as the "great-great-grandfather" of Google and the search engines that followed it.
W3Catalog and ALIWEB: The Web Tries to Index Itself
The web's earliest organizing attempts relied on people submitting their own content rather than anything automatically discovering it.
Oscar Nierstrasz's W3Catalog, released in September 1993, is generally credited as the web's first search engine. It worked by reformatting a set of existing, manually-maintained link lists into a single searchable format.
Weeks later, Martijn Koster built ALIWEB, short for Archie-Like Indexing of the Web, named directly after Emtage's earlier project. Like Archie and W3Catalog, ALIWEB worked by file submission rather than any discovery process like crawling. Site owners submitted their own page and a description, and ALIWEB indexed it. No bots, no automated discovery.
Few people, however, actually submitted. The same dependency that limited Archie's reach, requiring someone to register a source beforehand, shows up again later. Demonstrating that systems which are dependent on voluntary opt-in at scale will depend on how motivated people are to participate, and what value that participation actually brings.
The Yahoo Directory and the Bottleneck of Human Judgment
The next search approach added human judgment into the mix.
In January 1994, Stanford graduate students Jerry Yang and David Filo started "Jerry and David's Guide to the World Wide Web," which ultimately became the Yahoo Directory. This directory was a hierarchy of web pages organized into categories. Scripts automated the mechanical side, generating the directory's pages and running keyword search across its listings, but the filtering and categorization itself, deciding what a site was and where it belonged, stayed entirely manual. Filo specifically argued against automating it, on the grounds that no technology of the time could beat human judgment for that task.
At the time, this wasn't a temporary fix with a better plan in mind, but simply was the strongest option available. There weren't yet many pages to review, and automated search technology couldn't match a human editor's judgment about relevance. Yahoo's directory outperformed most automated competitors like InfoSeek and Excite specifically because manual review still had a quality advantage machines hadn't come close to.
However, the same human dependent constraint that limited ALIWEB and others was still present, but it was moved down the pipeline from submission/discovery to review.
600,000 Sites and a Directory That Can't Keep Up
Between 1993 and 1996, the number of sites on the web grew from roughly 130 to more than 600,000. Editorial staff size scaled linearly, but the level of growth didn't. These directory searches didn't fail quietly.
By 1998, the Yahoo Directory was still receiving 95 million page views a day. The model was sound, but it hit a ceiling that exponential growth eventually creates for anything bottlenecked by human labor.
In this same time window, a related development landed: NSFNET, the network we mentioned where commercialization was prohibited, was fully decommissioned and handed to private operators in April 1995. The restriction on commercial activity lifted at almost the same time the web's growth made human-curated discovery unsustainable.
JumpStation, WebCrawler, and AltaVista: Automation Takes Over
The structural fix arrived in stages.
JumpStation, built by Jonathon Fletcher in December 1993, was the first system to combine crawling, indexing, and search into one automated pipeline, indexing page titles and headers only. WebCrawler followed shortly after, indexing full page text, becoming the first full-text search engine. Both of which established the very discovery and indexing model every future search engine would use. A process where software discovers pages on its own and builds an index without requiring submission or curation.
Demand for this was immediate. When AltaVista opened its 16-million-document index to the public in December 1995, it received nearly 300,000 visits on its first day, with no marketing.
Discovery Is Solved. Now Who Pays?
By this point, automated search had solved the discovery problem, and the restriction on commercial activity had been lifted. That left an open question: who pays for this, and how?
Making a search engine free to use wasn’t the obvious answer, and in fact many directories charged for premium placement. Subscription access to online content was an established business model. AOL had built one of the largest companies of the era on it.
AOL's Walled Garden, and Why 34 Million Subscribers Wasn't Enough
Subscription models were tested at a real scale. By September 2002, AOL's subscription service had grown to its peak, roughly 26 million U.S. subscribers and over 30 million worldwide, the dominant consumer internet company of its era, built on a flat monthly fee for a bundle of internet access, email, and curated content.
The eventual decline of AOL traces back to a specific competition entering the field. Broadband providers like EarthLink and Comcast started selling plain internet access for less than AOL charged, with none of AOL's bundled software or portal attached. Once a household had that connection, the things AOL charged for, search, email, messaging, were already free elsewhere, so paying AOL on top of a connection for them stopped making sense to consumers.
AOL tried to compete on broadband directly, launching its own DSL offering through a partnership with Covad in 2004, but it repeated the same losing structure, where they charge consumers with a connection fee plus another for AOL's own software and features. The numbers kept moving one direction, and AOL lost 976,000 U.S. subscribers in a single quarter in 2006. By that summer it was down to 17.7 million subscribers, a 34 percent drop from its 2002 peak.
Why You Can't Sell a Search by the Query
Two separate factors explain why free won out over subscription specifically for search.
First, pricing: the value of a single search query is highly inconsistent and unknown before the search is performed. A query might be worth two cents or two hundred dollars depending on what's being searched for, and that's only known after the fact. There's no reliable way to set a per-query price under those conditions.
Second, growth mechanics: a subscriber's monthly fee doesn't improve the product for the next subscriber. A free user's attention and data make the product more valuable to advertisers, funding improvements that attract more free users. That loop reinforces itself in a way flat subscription pricing doesn't, which matters in a market where the largest player tends to take most of it.
GoTo.com and the Invention of Pay-Per-Click
Free was popular, but to stay profitable it had to address fundign problems. Problems that were eventually solved by the invention of pay-per-click advertisement models.
In February 1998, Bill Gross introduced GoTo.com at the TED conference. The pithc was that advertisers bid for placement in search results and paid only when a user actually clicked. This means that instead of paying for impressions and hoping it attracts a customer like a billboard on a highway, you only paid when you know your ad made the user act. This mattered specifically for search because a query expressed an explicit intent, something someone wants right now, which is more valuable to an advertiser than a passive view. Advertisers were paying up to a dollar per click within months.
Google's AdWords adopted the same mechanism a few years later and turned a profit in its first year running it.
Why Most Search Engines Since Have Been Free
Looking at the revenue models that were tried for search engines, only one really showed to survive. Human curation was no longer viable because it couldn’t match the growth. No directory or registry could keep pace with the growing size of the web, no matter how good its editors were. Charging by the query wasn’t an option because nobody can price it’s value until after it's used. AOL demonstrated that subscription models couldn’t survive in the market, even at the largest scale anyone could test it. All of this leaving free access as the only model still standing. Standing alone, however, only promised it to be popular, not profitable. Pay-per-click advertising is what cleared that last hurdle, turning growth into scalable profit. Leaving the ad model, the only survivor able to pay for itself.
GoTo Invented It. Google Won It.
Free, ad-funded search as a winning category doesn't end the story, though. GoTo, the company that invented pay-per-click, didn't end up running it. Yahoo, MSN, and AOL all eventually used some version of the same auction model GoTo built, and none of them became king of the hill either. The model itself wasn't the differentiator, since everyone running it had access to roughly the same mechanism. What separated the inventor from the eventual winner was a set of product questions the mechanism didn't answer on its own.
If the highest bidder always gets the spot, relevance becomes secondary to budget, and that's a real product risk since a results page optimized for who's willing to pay the most isn't necessarily optimized for what's actually useful. Therefore a search engine that stops being useful eventually stops getting used. Raising the question of not just how do you make money, but how do you make money in a way that doesn’t compromise the value the product has to consumers. GoTo and Google answered that question differently, and that difference decided who won. That's the subject of the next post: How Google Beat the Inventor of Search Advertising.