The Math of Machine Writing: How Computer Vision and NLP Flag Synthetic Text
We have been using various detection tools for a while now, to identify automated texts and AI-generated writing.
But have you ever imagined how these tools work to identify them? How do they differentiate between human language and sentences used by AI?
You have surely heard buzzwords like machine learning, LLMs, NLP, and computer vision. But do you know what they actually do to generate or detect synthetic and naturally human-written texts?
Since you are headed to an AI-powered future workspace where you must perform tasks using all these tools and platforms, you must know how they really work to help you, what byproducts they might have, and how you can avoid their negative effects.
Moreover, these are all the parts of AI literacy about which we must be aware to save our own privacy and avoid becoming dependent on them.
In this article, we will discuss what NLP and computer vision are and how they work to identify synthetic text.
What Are Computer Vision And NLP?
NLP, or natural language processing, is an amalgamation of the study of language and computational algorithms. NLP basically studies and analyzes AI and human language to help computers and tools differentiate between them.
It breaks down the different parts of linguistics, like syntax or sentence structure, grammar, vocabulary, and their arrangement, and the context of the whole text.
It helps chatbots and AI writing tools to read human language, understand their instructions, and respond accordingly.
Not only that, but it also analyzes human emotions and the mindset behind a particular instruction. So, they can give answers according to their moods and preferences.
The same mechanism is used in an AI checker to identify texts as human-written or AI-generated. Also, your email account automatically puts a particular email in spam using this NLP and machine learning.
You might wonder how these detection tools identify images or text in images, then. This is where computer vision(CV) comes in.
Similar to NLP, computer vision helps AI to read and interpret visual data. It works like human vision. Through this mechanism, AI is able to analyze the patterns and designs of an image.
This is the same mechanism that sensors and authenticators use for facial recognition for identity verification. It is also used across different industries like product quality control and medical diagnosis.
How Do They Flag Synthetic Text?
Both NLP and computer vision use some fundamental strategies and mathematical signals to flag synthetic or AI-generated texts. The most prominent of them are given below.
Perplexity
Remember solving math problems of probability?
You used to find out the chances of something likely to happen in percentage through a formula of probability. Unlike other math problems, that one was uncertain.
NLP uses a similar mathematical technique called perplexity or predictability. It analyzes the percentage of how much “surprising” a text is to a language model or a detection tool.
If a text appears to be more predictable and less surprising to a tool, that means it has low perplexity. And the tools end up detecting the text as AI-generated or synthetic.
On the contrary, if a text seems to be more surprising and unpredictable, it has high perplexity, which makes it a human-produced text.
Since human texts are likely to have more variety in language, showing ornamental vocabulary, they mostly have surprising elements and high perplexity. While AI writing tools generally use quite predictable words and patterns, they seem less surprising to AI, reflecting low perplexity.
So, this is how the tools and other AI detection platforms decide whether a text is human-written or AI-produced.
Analyzing Visual Data
For detecting images and text used in them, AI uses computer vision and optical character recognition (OCR) systems.
Through this tactic, it analyzes the visuals and data used, texts, patterns, and other artifacts used in them. Upon analysis, it extracts and identifies the text in the image, if there is any, and passes it to NLP. Then NLP does the rest to study the language.
To identify the image, whether it is synthetic or real, the OCR locates whether there are any unusual patterns in the image, or whether the living objects in it have overly smooth skin, plastic-like textures, and perfect physical features.
Also, computer vision measure if there are any structural inconsistencies, misaligned proportions, and stoicness or robotic patterns in the images.
When it comes to image identification, it is quite obvious whether it is real or automated because the theory is always very different from each other.
Burstiness
Burstiness is the variance in the syntactical structure, length, and use of wording across a whole writing piece.
NLP judges how many kinds of different sentence structures there are. Which kinds of words is the the writer has chosen and whether there is variety in them. Also, whether the sentences are of different lengths.
Usually, humans use an array of sentences with different structures, emotions, and purposes, which AI and chatbots do not. Instead, they remain fixated on certain sentence styles and words. So, these are the signals in the sentences that a platform or an NLP system uses to flag synthetic text.
Reading Historic Data
Lastly, AI stores, reads, and analyzes the historic data every time it needs to produce text or identify AI-generated portions in a text.
The conversations humans have with AI and the output the tools provide every time are stored as historical data. Whenever an AI-detection tool needs to detect a text, it finds the historical data and analyze them to decide.
If it finds similarities with the older AI-generated text, it flags the new one as synthetic text.
Structural Features
Another signal, which is kind of similar to burstiness, is reading the structural features.
If you see an AI-generated writing piece, you will see that the length of the sentences is mostly the same as well. Also, they follow some common sentence patterns and structure throughout the whole content, which makes it so obvious.
Moreover, they have a set of words, phrases, and connectors that they use mostly. They do not tend to go outside that fixation.
Finally, the texts will seem too grammatically accurate and perfect, which makes it sound robotic.
On the other hand, humans will showcase mistakes, variety of sentence patterna nd length, and more idioms and phrases. These pointers also help NLP to flag the text by an AI tool.
Final Thought
Though the detection tools have other identifiers and signals as well, which help them to declare a text as synthetic or natural, the pointers mentioned above are the fundamental ones.
From the rapid evolution and progress of LLMs and NLP, we can assume that they are going to be stronger and more effective in the future. So, AI detectors will be too.
But there will always be some differences between a human-written text and a synthetic text because humans always derive ideas from real-life experience, and chatbots mimic it.
