Skip to main content

Command Palette

Search for a command to run...

how Ai /LLM Work

Updated
•6 min read•View as Markdown

Artificial Intelligence has become part of our everyday lives. Whether we're asking ChatGPT to explain a programming concept, using an AI assistant to write an email, or getting coding help, We're already interacting with one of the most exciting technologies today—Large Language Models (LLMs).

  1. What is an LLM?

LLM stands for Large Language Model.

Let's break that down:

Large → Trained on enormous amounts of text.

Language → Designed to understand and generate human language.

Model → An AI system that has learned patterns from data.

Think of an LLM as someone who has read millions of books, articles, websites, research papers, and conversations. Instead of memorizing every sentence, it learns patterns in language—how words relate to each other and how ideas are expressed.

When you ask a question, it predicts the most suitable next words based on everything it has learned.

Popular Examples of LLMs

Some well-known LLMs include:

ChatGPT (OpenAI)

Gemini (Google)

Claude (Anthropic)

Llama (Meta)

DeepSeek

Mistral

Each model has its own strengths, but they all rely on similar core ideas.

Common Applications in Daily Life

You might already be using LLMs without realizing it.

Examples include:

Writing professional emails

Solving coding problems

Homework assistance

Creating social media captions

Customer support chatbots

Language translation

Resume writing

Interview preparation

Brainstorming business ideas

Content creation

Today, AI is becoming as common as search engines were a decade ago.

  1. What Happens When You Send a Message to ChatGPT?

When you type:

"Explain JavaScript promises."

A lot happens behind the scenes in just a few seconds.

Step 1: Typing a Prompt

our question is called a prompt.

It can be:

A question

An instruction

A paragraph

Code

Even a single word

Everything starts with your prompt.

Step 2: Processing Your Message

The AI doesn't immediately understand English.

Instead, it first prepares your text by breaking it into smaller pieces called tokens

Then these tokens are converted into numbers because computers only understand numerical data.

The model analyzes relationships between these numbers using a Transformer architecture.

Step 3: Generating a Response

The AI doesn't write the whole answer at once.

Instead, it predicts one token at a time.

For example:

Explain ↓

JavaScript ↓

Promises ↓

are ↓

used ↓

to ↓

handle ↓

...

Each new token depends on everything that came before it.

This happens incredibly fast until a complete response is generated.

Text vs Numbers

Computers perform calculations.

Everything inside a computer eventually becomes binary:

0 1 0 1 1 0

Whether it's:

Images

Videos

Audio

Documents

Text

Everything must become numbers first.

Why Convert Text into Numbers?

Imagine trying to calculate:

Apple + Banana

A computer can't perform math on words.

But it can calculate:

1024 + 584

Once words become numbers, AI models can compare them, identify patterns, and generate meaningful responses.

Introducing Tokens

Instead of processing entire paragraphs at once, AI breaks text into tokens.

Tokens are the basic pieces of text that an LLM works with.

We'll look at them next.

  1. Tokenization

Tokenization is one of the most important concepts in AI.

It means:

Breaking text into smaller units called tokens.

Think of it like cutting a pizza into slices before eating it.

The pizza stays the same, but smaller pieces are much easier to handle.

What Are Tokens?

A token might be:

A complete word

Part of a word

A number

A symbol

For example:

I love chai.

Could become:

"I"

"love"

"chai"

"."

Another word like:

unbelievable

might become:

un

believ

able

Different AI models use different tokenization methods.

Why Is Tokenization Needed?

Imagine reading a 500-page book.

Instead of trying to understand the entire book in one glance, you read it one word at a time.

LLMs work similarly.

Tokenization makes language easier to process efficiently.

Words vs Tokens

Many beginners think:

One word = One token

Not always.

Examples:

Text

Possible Tokens

Hello

Hello

unbelievable

un + believable

ChatGPT

Chat + GPT

2026

20 + 26

That's why token counts are often different from word counts.

What Is a Transformer?

A Transformer is a neural network architecture designed to understand relationships between words in a sentence.

Instead of reading one word at a time in strict order, it can examine the entire sentence and determine which words are most relevant to one another.

For example:

"The cat chased the mouse because it was fast."

A Transformer tries to figure out whether "it" refers to the cat or the mouse by looking at the surrounding context.

This ability to understand context makes it much better than many older approaches.

Why Did Transformers Change AI?

Earlier language models struggled with long sentences and complex relationships.

Transformers introduced mechanisms that made it easier to:

Handle long pieces of text.

Capture context more effectively.

Process information efficiently.

Scale to much larger datasets.

These improvements paved the way for today's powerful AI systems.

How Do Transformers Help Understand Language?

Instead of treating each word independently, a Transformer considers how every word relates to the others.

For example, in the sentence:

"The bank was crowded because people were depositing money."

The word bank clearly refers to a financial institution.

But in:

"We sat on the bank of the river."

The same word refers to the side of a river.

By analyzing context, Transformers can understand the intended meaning.

Why Do Almost All Modern LLMs Use Transformers?

Because they are:

Highly accurate

Scalable to billions of parameters

Excellent at understanding context

Efficient to train on massive datasets

Suitable for tasks like text generation, translation, summarization, coding, and question answering

That's why models such as ChatGPT, Gemini, Claude, Llama, and many others rely on Transformer-based architectures.

How Everything Works Together

Here's the complete high-level workflow of an LLM:

User │ ▼ Prompt │ ▼ Tokenization │ ▼ Convert Tokens into Numbers │ ▼ Transformer Processes Context │ ▼ Predict Next Token │ ▼ Generate Response │ ▼ Display Answer

From the user's perspective, this entire process feels almost instantaneous.

  1. Complete LLM Workflow

User │ ▼ Prompt │ ▼ Tokenization │ ▼ Embeddings │ ▼ Transformer Layers │ ▼ Predict Next Token │ ▼ Repeat Until Complete │ ▼ Final Response