Google published a groundbreaking research paper on identifying page quality using artificial intelligence. The details of the algorithm seem remarkably similar to what the useful content algorithm does.
Google Doesn’t Identify Algorithm Technologies
No one outside of Google can say for sure that this research paper is the basis of the useful content signal.
Google generally does not identify the underlying technology of its various algorithms such as the Penguin, Panda or SpamBrain algorithms.
So, it cannot be said with certainty that this algorithm is an algorithm of useful content, one can only speculate and offer an opinion about it.
But it’s worth a look because the similarities are eye-opening.
The Helpful Content Signal
1. It Improves a Classifier
Google has given many hints about the useful content signal, but there is still a lot of speculation about what it actually is.
The first clues were in a tweet dated December 6, 2022 that announced the first useful content update.
“It improves our classifier & working on global content in all languages.”
A classifier, in machine learning, is something that categorizes data (is it this or is it that?).
2. It’s Not a Manual or Spam Action
The Useful Content algorithm, according to Google’s explanation (What creators should know about Google’s August 2022 Useful Content Update), is not a spam action or a manual action.
“This classifier process is fully automated, using a machine learning model.
It’s not a manual action or an anti-spam action.”
3. It’s a Ranking Related Signal
The helpful content update explanation says that the helpful content algorithm is a signal used to rank content.
“…it’s just a new signal and one of many signals that Google evaluates to rank content.”
4. It Checks if Content is By People
Interestingly, the useful content signal (apparently) checks if the content is created by humans.
Google’s useful content update blog post (More content by people, for people in search) states that this is a signal to recognize content created by people and for people.
“…we’re making a number of improvements to Search to make it easier for people to find useful content made by and for people.
…We look forward to building on this work to make it even easier to find original content for real people in the coming months.”
The concept of content being “people” is repeated three times in the announcement, apparently indicating that this is a signal quality of useful content.
And if it wasn’t written by “humans,” then it was machine-generated, which is important to consider because the algorithm discussed here is related to the detection of machine-generated content.
5. Is the Helpful Content Signal Multiple Things?
Finally, Google’s blog post seems to indicate that updating useful content isn’t just one thing, like one algorithm.
Danny Sullivan writes that it’s “a series of improvements” which, if I’m not reading too much into it, means that it’s not just one algorithm or system, but several of them working together to remove the useless content.
“…we’re making a number of improvements to Search to make it easier for people to find useful content made by and for people.”
Text Generation Models Can Predict Page Quality
What this research paper reveals is that large-scale language models (LLMs) like GPT-2 can accurately identify low-quality content.
They used classifiers that were trained to recognize machine-generated text and found that these same classifiers could identify low-quality text, even though they were not trained to do so.
Great language models can learn how to do new things they were not trained to do.
A Stanford University article on GPT-3 talks about how it learned to translate text from English to French on its own, simply because it was given more data to learn from, something that didn’t happen with GPT-2, which was trained on less data.
The article states that adding more data causes new behaviors to emerge, a result of what is called unsupervised training.
Unsupervised training is when a machine learns how to do something it was not trained to do.
That word “emerge” is important because it refers to when a machine learns to do something it was not trained to do.
A Stanford University article on GPT-3 explains:
“Workshop participants said they were surprised that such behavior emerged from simply scaling data and computing resources, and expressed curiosity about what further possibilities would emerge from further scaling.”
A new emerging capability is exactly what the research paper describes. They found that a machine-generated text detector can also predict low-quality content.
“Our work is twofold: first, we show by human evaluation that classifiers trained to distinguish between human and machine-generated text emerge as unsupervised predictors of ‘page quality’, capable of detecting low-quality content without any training.
This allows quality metrics to be run quickly in low-resource settings.
Second, curious to understand the prevalence and nature of low-quality pages in the wild, we are conducting an extensive qualitative and quantitative analysis of over 500 million web articles, making this the largest study ever conducted on the topic.”
The conclusion is that they used a text generation model trained to spot machine-generated content and discovered that a new behavior appeared, the ability to recognize low-quality pages.
OpenAI GPT-2 Detector
The researchers tested two systems to see how well they performed at detecting low-quality content.
One of the systems used RoBERT, which is a pre-training method that is an improved version of BERT.
These are the two tested systems:
They found that OpenAI’s GPT-2 detector was superior in detecting low-quality content.
The description of the test results closely reflects what we know about the payload signal.
AI Detects All Forms of Language Spam
The research paper states that there are many signals of quality, but that this approach focuses only on language or linguistic quality.
For the purposes of this algorithm research paper, the phrases “page quality” and “language quality” mean the same thing.
The breakthrough in this research is that they successfully used the OpenAI GPT-2 detector’s prediction of whether something is machine-generated or not as a result of language quality.
“…documents with a high P (machine written) score tend to have low language quality.
…Machine authorship detection can therefore be a powerful proxy for quality assessment.
It doesn’t require any annotated examples – just a body of text that you can practice on in a self-judgmental way.
This is particularly valuable in applications where labeled data are sparse or where the distribution is too skewed for good sampling.
For example, it is challenging to organize an annotated dataset representative of all forms of low-quality web content.”
This means that this system does not need to be trained to detect certain types of low-quality content.
It only learns to find all low quality variants.
This is a powerful approach to identifying pages that are not of high quality.
Results Mirror Helpful Content Update
They tested this system on half a billion web pages, analyzing the pages using various attributes such as document length, content age and topic.
Age of content does not mean new content is marked as low quality.
They simply analyzed web content by time and found that there was a big jump in low-quality pages starting in 2019, which coincided with the growing popularity of using machine-generated content.
Analysis by topic revealed that certain topic areas tend to have higher quality pages, such as legal and government topics.
Interestingly, they found a large amount of low-quality sites in the education space, which they said corresponded to sites that offered essays to students.
What makes this interesting is that education is a topic that Google specifically mentioned will be affected by the useful content update.
A Google blog post written by Danny Sullivan shares:
Three Language Quality Scores
“…our testing has shown that it will especially improve outcomes related to online education…”
Google’s Quality Assessor Guidelines (PDF) use four quality scores, low, medium, high, and very high.
The researchers used three quality scores to test the new system, plus another called undefined.
Documents rated as undefined were those that could not be rated for any reason and were removed.
Grades are graded 0, 1 and 2, with two being the highest grade.
These are the descriptions of the language quality (LQ) scores:
“0: Low LQ.
The text is incomprehensible or logically inconsistent.
1: Medium LQ.
Lowest Quality:
The text is comprehensible, but poorly written (frequent grammatical/syntactic errors).
2: High LQ.
The text is understandable and relatively well written (rare grammar/syntax errors).
Here are the definitions of low quality in the Quality Assessor Guidelines:
“MC was created without the appropriate effort, originality, talent or skill required to satisfactorily achieve the purpose of the site.
…little attention to important aspects such as clarity or organization.
… Some low-quality content is created with little effort to have content to support it
monetization instead of creating original or hard content to help users.
Filler” content can also be added, especially at the top of the page, forcing users to scroll down to get to the MC.
…The writing of this article is unprofessional, including many grammatical and punctuation errors.”
The guidelines for quality assessors have a more detailed description of low quality than the algorithm.
The Algorithm is “Powerful”
It is interesting how the algorithm relies on grammatical and syntactic errors.
Syntax is a reference to word order.
Words in the wrong order sound incorrect, similar to what the character Yoda says in Star Wars (“Impossible to see the future is”).
Does the useful content algorithm rely on grammar and syntax signals? If this is an algorithm, then maybe that can play a role (but not the only role).
But I’d like to think that the algorithm has improved with some of what’s in the QA guidelines between the release of the research in 2021 and the introduction of the useful content signal in 2022.
It’s good to read what the conclusions are to get an idea of whether the algorithm is good enough to use in search results.
Many research papers end by saying more research is needed or conclude that improvements are marginal.
The most interesting works are those that seek new results.
The researchers note that this algorithm is powerful and outperforms baseline values.
What makes this algorithm a good candidate for a useful content type signal is that it is a low-resource algorithm that is on web metrics.
In conclusion, they reaffirm the positive results:
“This paper argues that detectors trained to distinguish human from machine-written text are effective predictors of web page language quality, outperforming a basic supervised spam classifier.”
The conclusion of the research paper was positive regarding the findings and the hope was expressed that the research will be used by others.
Citations
Google Research Page:
No mention is made that further research is needed.
Download the Google Research Paper
This research paper describes advances in low-quality web page detection.
The conclusion shows that, in my opinion, there is a probability that it could enter the Google algorithm.
What algorithm is used in search?
Since it’s described as a “web-scale” algorithm that can be deployed in a “low-resource setting,” that means it’s the kind of algorithm that could be active and run on a continuous basis, just like the useful content signal is said to do .
We don’t know if it’s related to a useful content update, but it’s certainly a breakthrough in the science of low-quality content detection.
What are the 2 types of searching algorithms?
Generative models are unsupervised predictors of page quality: a study of colossal scale
What are the two most common search algorithms?
Generative models are unsupervised predictors of page quality: a colossal scale study (PDF)
What are the types of searching algorithm?
Featured Image Shutterstock/Asier Romero
- Linear Search Algorithm Linear search algorithms are considered the most basic of all search algorithms because they require a minimal amount of code to implement. Also known as sequential search, linear search algorithms are the simplest formula for using search algorithms.
- What algorithm is used in SEO? What is the Google algorithm for SEO? As mentioned earlier, the Google algorithm uses keywords in part to determine page rankings. The best way to rank for certain keywords is SEO. SEO is essentially a way of telling Google that a website or website is about a certain topic.
- There are two types of search: sequential search and interval search. Almost every search algorithm falls into one of these two categories. Linear and binary search are two simple algorithms that are easy to implement, and binary algorithms run faster than linear ones.
- Roughly speaking, there are two categories of search algorithms that you need to know right away: linear and binary.
- Search algorithms:
- Linear search.
- Binary search.
- Jump Search.
Why does Google keep its algorithm secret?
Interpolation search.
Exponential search.
How is Google algorithm protected?
Search sublists (search a linked list in another list)
How do I control Google algorithm?
Fibonacci search.
- Ubiquitous binary search.
- Google’s response Google has made it clear in the past that it will not disclose its algorithm for two main reasons: The algorithm is a trade secret. Its disclosure would give the competition an advantage. Revealing the algorithm would be an invitation to all the spammers in the world, resulting in a vastly inferior web.
- What is the purpose of Google’s algorithm? Google’s algorithms are a complex system used to retrieve data from the search index and instantly provide the best possible results for a query. A search engine uses a combination of algorithms and a number of ranking factors to deliver web pages ranked by relevance on its search engine results pages (SERPs).
- Like newspapers, Google’s algorithms are protected by the First Amendment, making them difficult to regulate by law.
- How to succeed with the Google algorithm
- Optimize for mobile devices. …
- Check your incoming links. …
- Increase user engagement. …
Is Google search algorithm patented?
Reduce page load time. …
Is Google search algorithm patented?
Avoid duplicate content. …
Which technique did Google get a patent for?
Create informative content. …
Is Google’s search algorithm a secret?
Avoid keyword stuffing. …
Is Google algorithm public?
Don’t over-optimize.
What is Google’s algorithm?
Lawrence Page, one of the co-founders (with Sergey Brin) of Google, developed the PageRank algorithm in 1997. On January 9, 1998, Brin filed a patent application, and on September 4, 2001, the patent was granted.
Is the Google algorithm secret?
Lawrence Page, one of the co-founders (with Sergey Brin) of Google, developed the PageRank algorithm in 1997. On January 9, 1998, Brin filed a patent application, and on September 4, 2001, the patent was granted.
How does Instagram algorithm work?
On June 21, 2012, Google was granted a patent to record information in a query log file that tracks your click data and ranks web pages based on clicks coming from specific locations.
And although Google provides SEO insiders with frequent updates, the company’s search algorithms are a black box (a trade secret it doesn’t want to give away to competitors or spammers who will use it to manipulate the product), meaning that knowing what kinds of information Google will privilege is demanding. …
How do you beat the algorithm on Instagram?
Google’s algorithm is extremely complex, and exactly how it works is not public information. There are believed to be more than 200 ranking factors, and no one knows them all. Even if it does, it won’t matter because the algorithm is always changing.
- What is the Google algorithm? The Google search algorithm is a complex system that allows Google to find, rank, and return the most relevant pages for a given search query. To be precise, the entire ranking system consists of multiple algorithms that take into account various factors such as the quality, relevance or usability of the page.
- Within Google’s search algorithm, understanding and clarifying the meaning and intent of a search query is a crucial first step. The mechanisms that make this possible are, again, a mystery, but we do know that it allows the search engine to understand: the scope of the query.
- The Instagram algorithm is a set of rules that rank content on the platform. It decides what content appears and in what order on all Instagram users’ feeds, the Explore page, the Reels feed, hashtag pages, etc. The Instagram algorithm analyzes every piece of content published on the platform.
- What are the 3 main factors of the Instagram algorithm? The Instagram algorithm is divided into 3 main factors: INTEREST, RELATIONSHIP, TIMELINESS.
- 16 Ways to Beat the Instagram Algorithm in 2022
- Create reels often and consistently.
- Keep up with trends.
- Make stories fun.
Why is my Instagram not reaching anyone?
Post stories daily.
Why is my reach so low on Instagram 2022?
Use hashtags wisely.
What triggers Instagram algorithm?
Get your content shared.
What affects the Instagram algorithm?
Carousel post instead of one photo.
How does Instagram decide what to show you?
Write good subtitles.
How does Instagram 2022 algorithm work?
If you’re not posting regularly, you’ll likely notice a drop in reach. Consistent does not mean multiple times a day, every day. In fact, unless you have a team to help you do it, you’re bound to burn out. Consistency is one of the most talked about elements of success on Instagram, but it’s also the hardest.
How does the Instagram algorithm work now?
In the last few years, I have noticed a significant drop in impressions, engagement and the number of new followers. The reason: Instagram’s algorithm in 2022 Thanks to the last few updates, only 10% of your followers can see your post.
Why is my reach so low on Instagram 2022?
Instagram uses a combination of algorithms, processes and classifiers to determine the most relevant content for each user. It considers various signals such as user activity, post information, interaction history and content creator information.
What’s an algorithm used in social media?
The algorithm takes into account hundreds of factors such as user history, location, profile, device, trends, relevance, popularity, etc. From SEO to social media, algorithms are often the ones that determine who actually sees the content you post and who doesn’t.
You may see suggested posts in places like your Instagram feed and Explore. These suggestions are based on things like: Your activity: Who you follow and which posts you’ve liked, saved, or commented on. Your Connections: Your connection history with that account or similar accounts on Instagram.
What is the algorithm of Facebook?
A new Instagram algorithm dictates the order of posts that users see as they scroll through their feed. Based on specific signals, it prioritizes the best posts, pushing the most relevant ones to the top and giving them the most visibility, while other content ends up being placed lower.
How do you beat Facebook algorithm?
Every time a user opens the app, Instagram algorithms instantly comb through all the available content and decide which content to serve him (and in what order). The 3 most important ranking factors of the Instagram algorithm in 2022 are: The relationship between content authors and viewers.
- In the last few years, I have noticed a significant drop in impressions, engagement and the number of new followers. The reason: Instagram’s algorithm in 2022 Thanks to the last few updates, only 10% of your followers can see your post.
- A social media algorithm is a set of rules and signals that automatically rank content on a social platform based on how likely each individual social media user is to like and interact with it.
- How do social media algorithms work in 2022? Social media algorithms work using a set of rules and data to determine which posts to show to each individual social media user. The goal is to organize very interesting feeds that will keep people active on the platform as much as possible.
- Facebook’s algorithm is a set of rules that decide which posts people see in their Feeds. In essence, it decides which content is most relevant to display to each user based on several factors. Each user’s feed will look very different because it is personalized just for them.
- Strategies for using the Facebook algorithm to your advantage
- Create and share great content. …
- Generate user interactions. …
- Answer, answer, answer. …
How does Facebook algorithm work now?
Jump into the (live) video feed. …
What are examples of algorithms?
Consider Facebook ads. …
What is a algorithm give an example?
Go local. …
What are 5 examples of algorithms?
Get your team involved. …
- Ask your fans to follow you.
- Facebook’s algorithm determines which posts your audience sees in their feed based on their interactions and behavior. The algorithm takes into account inventory, signals, predictions and relevance score on Facebook to decide which posts will rank high in 2022.
- Common examples include: a recipe for baking a cake, the method we use to solve the long division problem, the laundry process, and the functionality of a search engine are all examples of algorithms.
- What is an algorithm? An algorithm is a set of instructions for solving a problem or performing a task. One common example of an algorithm is a recipe that consists of specific instructions for preparing a dish or meal.
- Examples of algorithms in everyday life
- Shoe tying.
- The following recipe.
What algorithms are used in social media?
Classification of objects.
- Bedtime routines.
- Finding a book from the library in the library.
- Driving to or from Somewhere.
- Deciding what to eat.
Which machine learning algorithm is used for social media?
Types of Social Media Algorithms
What algorithm is Facebook?
Popularity.
How many algorithms does Google use?
Content type.
Relationship.
How big is Google’s algorithm?
Recency.
How much is the Google algorithm worth?
Chatbot System Whether it’s e-commerce or social media, chatbot systems are a well-known application of machine learning.
How many algorithms does Google use?
User Affinity: Part of the user affinity algorithm in Facebook’s EdgeRank looks at the relationship and proximity of users to content (post/status updates).
How many types of Google algorithms are there?
You may already know that Google uses over 200 ranking factors in its algorithm… But what exactly are they?
What are the types of Google algorithm?
What algorithm does Google use? PageRank (PR) is an algorithm used by Google Search to rank web pages in search engine results.
- Google’s algorithm is extremely complex, and exactly how it works is not public information. There are believed to be more than 200 ranking factors, and no one knows them all. Even if it does, it won’t matter because the algorithm is always changing.
- How much is the algorithm worth? Hundreds of billions, potentially. Ask Google (NASDAQ:GOOG) . Its search algorithm was worth $180 billion as of yesterday’s close.
- You may already know that Google uses more than 200 ranking factors in its algorithm… But what exactly are they?
- There are 9 types of Google algorithms: Panda: Google Panda is a major change to Google’s search ranking algorithm that was first announced on February 24, 2011.
- Google’s algorithms are a complex system used to retrieve data from the search index and instantly provide the best possible results for a query… What are Google Algorithms?
- Florida.
- Big daddy.
- Jagger.
What is Google’s search algorithm?
Vince.
How many times does Google change algorithm?
Caffeine.
When was the last Google algorithm update?
Panda.
What is Google’s latest search algorithm?
Freshness algorithm.
