The Quiet Ai.
← Articles

My Thai Teacher and My Thai Classroom Don’t Have to Agree

A voice teacher and a separate classroom disagreed about one answer in Lesson 04. The classroom was right.

A loop of eight steps (learner state, lesson authoring, design, classroom build, voice teaching, evidence, assessment, research), with steps 1, 5, 6 and 7 highlighted and two separate panels inside it. The left panel sets “The voice teacher told me I was right.” against “The classroom said I was wrong.” across a not-equal sign, above “The classroom was right.” The right panel starts from “The teacher’s report said I was beginning to transfer one of the sentence structures into new situations”, then “The ICM compares all of it. No. That sentence had already appeared earlier in the lesson.”, then “The teacher had overstated what happened.”, ending in “insufficient evidence.”

I’ve been building a Thai language classroom around a strange idea:

What if the AI teacher and the classroom were two different systems?

One talks to me.

One shows me things.

And neither one has to do everything.

I’ve now run a baseline conversation and three real lessons through it, and Lesson 04 changed how I think about the whole thing.

Because this time, the classroom caught the teacher being wrong.

And the teacher was AI.

Well... ChatGPT acting as my Thai teacher.

Here’s the setup

I have ChatGPT Voice on my phone.

That is my teacher.

It listens to my Thai, talks back to me, corrects me, asks questions, hears when I hesitate, and can change direction when something isn’t working.

On my computer is a completely separate HTML classroom.

It doesn’t have a giant language model deciding what to show next.

It renders a lesson.

Images. Thai text. Sentence structures. Reading. Multiple choice. Ordering. Vocabulary. Visual explanations.

And underneath all of this is an ICM structure keeping track of what we’ve learned about me as the student.

That structure is built on Interpretable Context Methodology (ICM), Jake Van Clief and David McDermott’s way of organizing an AI’s work as folders and files.

The interesting part is that the voice teacher and classroom aren’t directly connected.

They share the same lesson.

The lesson is the contract between them.

A phone showing a sound wave, captioned “One talks to me.”, and a laptop showing a painted scene of a man carrying a surfboard through rain, captioned “One shows me things.” A card headed “Lesson 04”, with a line of Thai text on it, is linked by a line to both devices, under the heading “The lesson is the contract between them.”

Lesson 01 was basically: “Who is this student?”

We talked.

I spoke Thai.

I told it something important, and the baseline confirmed it.

My problem isn’t that I can’t communicate in Thai.

I can.

I can tell stories, explain things, talk about experiences and usually get where I need to go.

But I often build Thai sentences as if I were constructing them in English.

That gave the system somewhere to start.

Lesson 02 was too easy.

The system built a lesson around:

ถ้า...ก็...

“If...then...”

And I basically flew through it.

That was useful.

Not because Lesson 02 was great.

Because it produced evidence that I’d been underestimated.

So Lesson 03 changed.

More reading.
More sentence comparison.
More visual material.
More complicated relationships between ideas.

It was considerably better.

But it exposed another problem.

We had built exercises.

We hadn’t really built a class.

That distinction sounds small.

It wasn’t.

The classroom was getting very good at asking me things:
Which sentence is correct?
Put these pieces in order.
Read this.
What does this mean?
Speak about this picture.

But at one point I realized:

When did anybody actually teach me this?

We had accidentally built a sophisticated testing machine.

So before Lesson 04, we changed the lesson structure.

Now an activity could explicitly be:

TEACH
GUIDED PRACTICE
INDEPENDENT PRACTICE
TRANSFER

And TEACH meant exactly that.

Don’t test me.

Teach me something.

Four steps rising left to right like stairs: 1 TEACH (“Don’t test me. Teach me something.”), 2 GUIDED PRACTICE (“Put these pieces in order.”), 3 INDEPENDENT PRACTICE (“Then less help.”), 4 TRANSFER (“Something I haven’t seen.”). The heading reads “When did anybody actually teach me this?” and a bracket under the last three steps reads “We had accidentally built a sophisticated testing machine.” In the animated version the steps light up one at a time.

Lesson 04 started feeling like a real classroom.

There were visual explanations.
Sentence anatomy.
Contrasts.
A reading section.
Vocabulary.
Images.
Then guided practice.
Then less help.

Then independent work.

Meanwhile ChatGPT Voice was beside me teaching the lesson.

And this is where things became interesting.

Because the teacher made mistakes.

At one point it misunderstood what was on my screen.

At another point it jumped in too quickly while I was still trying to construct an answer.

It even randomly switched languages on me. I had to stop it:

“คุณพูดภาษาจีน เราเรียนภาษาไทย.”

You’re speaking Chinese. We’re learning Thai.

But the really interesting mistake happened when I classified a sentence incorrectly.

The voice teacher told me I was right.

The classroom said I was wrong.

The classroom was right.

That moment changed something for me.

The systems don’t need to agree.

Originally, I was thinking mostly about coordination.

How do I make the visual classroom and voice teacher work together?

But Lesson 04 showed me something better.

I don’t necessarily want them to agree.

They have different jobs.

The voice teacher can make a linguistic judgment in the moment.

The classroom can record what happened on screen.

And afterward, the ICM can compare:

What did the lesson intend?
What did Gabe actually do?
What did the classroom record?
What did the teacher think happened?

Those aren’t always the same thing.

And after Lesson 04, they weren’t.

The teacher’s report said I was beginning to transfer one of the sentence structures into new situations.

The assessment went back through the lesson and essentially said:

No.

That sentence had already appeared earlier in the lesson.

I had successfully used it again.

But that wasn’t evidence of spontaneous transfer.

The teacher had overstated what happened.

So the assessment changed the result to:

insufficient evidence.

I loved that.

Because now this wasn’t just AI remembering what I did.

One part of the system was capable of challenging another part’s interpretation of what I did.

So we changed the workflow again.

This is roughly what happens now:

1. Learner state
Before the lesson, the system has a record of what I appear to know, what remains uncertain, what I’ve struggled with, and what hasn’t actually been tested yet.

2. Lesson authoring
The teacher designs the next learning experience from that state.

3. Design
The Designer doesn’t just “make it pretty.”
Its job is to ask:
Can a visual make this idea easier to understand?
Sometimes that means an illustrated scene.
Sometimes sentence anatomy.
Sometimes a diagram.
Sometimes almost nothing.

4. Classroom build
The lesson becomes the deterministic HTML classroom on my computer.

5. Voice teaching
ChatGPT Voice teaches beside it.
The classroom handles the visual and interactive world.
The teacher handles conversation, pronunciation, explanation and adaptation.

6. Evidence
The classroom records what happened on screen.
The teacher writes its own report.
I can add my experience as the learner.

7. Assessment
The ICM compares all of it.
And importantly:
it doesn’t have to believe the teacher.

8. Research
This time, before Lesson 05, we added another worker.
A Researcher.
Instead of us simply saying:
“PowerPoint-style teaching seems better.”
we asked:

What does the research actually say about teaching adults from screens?
What about second-language learners?
Should text remain on screen while a teacher reads it?
Should English translations always be visible?
How much grammar explanation is useful?
When should scaffolding disappear?
Should the next lesson begin with a recap or retrieval?

The Researcher came back with something I particularly liked:

It disagreed with some of our assumptions.

Some ideas were strongly supported.

Some had mixed evidence.

Some things we’d been treating as “good teaching” turned out to be mostly teaching convention.

And two of our own earlier suggestions didn’t survive it. One was dropped, one had to be rewritten.

That’s exactly what I want from a research worker.

Not:

“Find research proving Gabe’s idea.”

But:

“Tell us whether Gabe’s idea survives the research.”

Now we’re building Lesson 05 differently.

And here’s another change.

I told the classroom:

Enough rain and surfing.

Because once a system finds useful context about you, there’s another danger.

It can become trapped by its own personalization.

Gabe surfs.

Gabe teaches yoga.

Gabe lives in Thailand.

Suddenly every Thai sentence involves rain, surfing and yoga.

That’s not personalization anymore.

That’s a very small universe.

So Lesson 05 is going somewhere else.

I want more complicated sentence structures.

More reading.
Real Thai from the world around me.
Street signs.
Notices.
Menus.
Messages.
The more complex signs you see driving around Thailand.

Even the prayer and chanting boards inside Thai temples.

Not because “authentic material” sounds impressive.

Because I want to walk out of the lesson and recognize something I couldn’t read before.

On the left, four painted lesson illustrations in a circle (a man carrying a surfboard in the rain, a man at a window, a man looking at the sea), captioned “Enough rain and surfing.” and “That’s a very small universe.” On the right, under “Real Thai from the world around me.”, four photographs: a blue notice board in Thai (tagged “Notices.”), a temple gate, a roadside sign (tagged “Street signs.”) and a park information board in Thai and English.

And there’s another lesson we’ve learned:

Lesson 05 cannot feel like an assessment of Lesson 04.

The system absolutely needs to find out whether I retained what we worked on.

But that’s its problem.

Not mine.

As the learner, I should walk into:

a new Thai class.

Something interesting.

Something I haven’t seen.

Something worth learning.

And somewhere inside that experience, the system can quietly discover whether yesterday’s learning survived.

That’s becoming one of my favorite principles from this experiment:

The learner experiences the lesson.
The system measures the learning.

Those should not feel like the same thing.

Two halves. On top, “The learner experiences the lesson.” beside a lesson card listing four Thai sentence patterns with English explanations. Below, “The system measures the learning.” beside an abstract panel of grey lines with one row highlighted in amber and no readable data; a pill between the halves reads “Those should not feel like the same thing.”

Four sessions ago, this was basically:

Can I make ChatGPT Voice and a webpage teach me Thai at the same time?

Now there’s a learner model, lesson authoring, a Designer, a deterministic classroom, a voice teacher, evidence coming from different sources, an assessment layer capable of disagreeing with the teacher, and a Researcher whose findings feed the next lesson.

And somehow, underneath all of that complexity, my experience is becoming simpler.

I sit down.

I open the classroom.

I turn on my teacher.

And I learn Thai.

Lesson 05 is next.

And yes.

This time I’m filming it. 😁

Less hype, more attention.

If this is the kind of slow, unglamorous, actually-works thinking you want more of, that’s the conversation I have most days.

Work with me →

Or keep reading — more articles.