Sponsors

Stay updated with our latest job postings by following us on LinkedIn and join our Discord community for daily notifications.

Recap of Qwen Meetup Seoul #4, where 130 developers explored cheaper AI coding with Qwen and how TiDB keeps production AI agents reliable.

Qwen Meetup Seoul #4: AI coding costs and the reality of production agents

이 글은 한국어로도 제공됩니다.한국어로 읽기 →

Want to feel what the night was like? From the talks to the networking and everything in between, here's the aftermovie from Qwen Meetup Seoul #4 🎥.

On August 19 we ran Qwen Meetup Seoul #4 together with Qwen at dcamp Seolleung. Around 130 people came out for the evening.

Both talks were about problems that become a lot more obvious once you start using AI seriously: first, how expensive coding agents can get when you're running frontier models for everything, and second, what happens when the AI agent you built stops being a demo on your laptop and has to survive in production.

Different layers of the stack, but actually a pretty similar question underneath both:

How do you make AI systems efficient enough to keep using them at scale?

Wide shot of the room

You probably don't need your most expensive model for everything

Edwin Tack, Solutions Architect at Alibaba Cloud Korea, opened with a question for the room: what are people actually using for coding every day?

A lot of hands went up for Claude Code. Some for Codex. Others were using models like DeepSeek.

His starting point was simple. Model performance matters, obviously, but once every engineer on a team starts running agents all day, the cost starts mattering too.

And maybe the answer isn't to choose one model.

Edwin Tack presenting

Edwin split an AI coding workflow into two broad stages:

Planning and design, where you want the strongest reasoning you can get.

And execution, where the agent is spending much more of its time actually writing, modifying and iterating on code.

The proposal was to keep a frontier model for the planning stage, but use something cheaper such as Qwen3.8 Max for execution.

Instead of asking which model is best, ask which model is best for this particular step.

According to the estimates Edwin presented, a hybrid setup could cut model costs by roughly 30 percent, while moving both planning and execution to Qwen could push the reduction closer to 50 percent.

[Photo: slide showing the hybrid model / cost comparison]

I liked the analogy he used later in the talk.

Infrastructure teams rarely assume one cloud provider must run absolutely everything. Having AWS, Azure, GCP, Alibaba Cloud or another provider competing for different workloads gives companies options.

His argument was that AI models are heading in the same direction.

We shouldn't necessarily think of Claude, GPT, Qwen, GLM, Kimi or DeepSeek as one permanent choice. The application can route work between them depending on quality, latency and cost.

That becomes even more interesting as the performance gap between models gets smaller.

The model gap is shrinking

The Q&A went quite deep into this.

One question was basically: cost is nice, but time is money too. If Qwen takes longer to produce the same result, are you actually saving anything?

Edwin's view was that the gap between frontier models and the models following them is getting increasingly small, especially for many coding workloads.

New models keep arriving, benchmarks keep moving, and something that looks comfortably ahead today might have several close competitors a few months later.

His prediction was that cost will therefore become a bigger part of the decision.

Another attendee asked what happens to AI pricing over the next few years. Edwin expects the most powerful frontier models to remain expensive, while cheaper competitors put pressure on how far those prices can actually rise.

That leads back to the hybrid approach: spend the expensive tokens where they make the biggest difference, and don't spend them where they don't.

Audience member asking a question

There were also questions about running smaller Qwen models locally, passing context between models, and how to plug this kind of setup into existing coding tools.

The nice part was that the discussion wasn't really "Qwen versus everything else."

It was much more practical:

If I already use the models I like, where else in my workflow could another model make sense?

Link: How to optimize Vibe coding costs with Qwen3.8 Max — full talk

Your agent isn't a chatbot

Then PD, Senior Solution Architect at TiDB, took the stage and started somewhere I definitely wasn't expecting: with a camera he lost in the sea in Thailand 17 years ago.

He had been climbing back onto a small boat after snorkeling when the camera hanging around his neck got caught and dropped into the water.

The camera itself wasn't what bothered him.

The photos were.

A month later, somehow, a diver found it around ten meters underwater. Someone posted the recovered photos online, one of PD's friends recognized him, and he eventually got the camera and his memories back.

He still has those photos today.

PD showing the old camera / Thailand story

That was his setup for the rest of the talk.

The valuable part of an AI agent often isn't the process itself.

It's the state the process accumulated while working.

And unlike a simple chatbot, an agent might have a whole sequence of work to complete:

fetch the order, check the policy, calculate something, call another API, issue the refund, update the result.

Now imagine the process dies halfway through.

If all of that state lived only inside the running process, the new process doesn't know what already happened.

It starts again.

When restarting means paying twice

The harmless version of that problem is wasted tokens.

An agent gets through six steps, the pod dies, it restarts from step one, and you pay again for work it already completed.

PD demonstrated exactly that.

In his example, restarting the workflow from scratch used roughly 2,300 tokens in total. Once the agent wrote a checkpoint into TiDB after each completed step, it could recover from the last successful checkpoint instead.

The same interrupted workflow finished at roughly 1,300 tokens.

Checkpointing demo / terminal

But tokens aren't the scary part.

Imagine the action halfway through the workflow was a refund.

The agent checks that the customer qualifies, calculates the amount, sends the refund, and then crashes before remembering that the refund was sent.

It starts over.

Congratulations to the customer. They may now have two refunds.

PD's point was that we've spent decades learning how to make financial and transactional systems reliable, and then started building agents that keep important state inside a Python process.

Putting an LLM in the loop doesn't remove the old distributed systems problems.

It adds new ways to trigger them.

The prototype worked. Then users arrived.

This was probably the most useful distinction in the talk.

Building an agent prototype today is surprisingly easy.

You have your model, maybe PostgreSQL or MySQL for application data, a vector database for embeddings, Redis for some state, and APIs connecting everything.

It works.

You demo it after a weekend and look like a genius.

Then you get real traffic.

Processes run out of memory. Pods restart. Deployments replace instances while agents are halfway through a task. Hundreds or thousands of workflows start writing state concurrently.

Meanwhile your transactional database has one version of the data, your vector database has another, and the embedding pipeline hasn't caught up with yesterday's policy change yet.

Those are problems you don't see when ten developers are testing the prototype.

You see them when customers arrive.

PD described TiDB's role as a persistent layer underneath the agents: keep checkpoints in a normal SQL database, recover workflows after failures, and handle the concurrency once the application grows past the point where a single database node is comfortable.

The larger architecture he showed combines transactional data, vector search, full-text search and analytical workloads rather than maintaining a separate database and synchronization pipeline for every one of them.

Yes, you can do checkpointing with PostgreSQL

Someone eventually asked the question PD knew was coming:

Couldn't you just do this with PostgreSQL?

His answer was yes.

The checkpoint itself isn't magic. You can write a row containing the task ID, completed step and relevant state into almost any database.

His argument for TiDB starts when the workload becomes large.

For a small application, he was very clear that PostgreSQL can actually be faster because the architecture is simpler.

But once concurrency keeps increasing, one primary writer eventually becomes a bottleneck. TiDB distributes writes across nodes and can add capacity horizontally without requiring the application to be redesigned around a new database architecture.

That distinction came up several times during the Q&A:

TiDB isn't supposed to win the tiny benchmark. It's supposed to keep behaving predictably after the tiny benchmark stops resembling your production system.

There were good follow-up questions on Raft consensus, write latency, strong versus eventual consistency, storage tiering, embedding models, and how TiDB can serve both transactional and analytical workloads.

Second Q&A session

Link: When AI apps become real: The hidden infrastructure behind agents in production — full talk

Two different talks, the same constraint

Looking back, the talks connected better than I expected.

Edwin's talk was about not wasting your most expensive model on work that a cheaper model can do.

PD's talk was about not wasting tokens doing work an agent already completed.

One optimizes which model performs the work.

The other makes sure you don't perform the same work twice.

As agents become a normal part of software products, those details matter more.

A demo can ignore a few extra dollars in tokens. It can ignore the occasional restart. It can use one model for everything and keep state wherever is easiest.

A production system with thousands of users can't.

People talking after the sessions

The room

Around 130 people joined us this time.

We started with food and check-in, moved into the two talks and Q&As, and then kept the room open for networking afterward.

One thing I've been enjoying about these AI meetups is how technical the conversations continue to get after the official program ends. You hear people comparing models, architectures, tools and things they're actually building rather than just talking about AI at a high level.

That's what I want these events to be.

Networking

Huge thanks

This event wouldn't have happened without:

  • Edwin Tack for breaking down where a multi-model coding workflow can actually save money
  • PD for turning a lost camera into a surprisingly good explanation of persistent agent state
  • Soo & Febria for continuing to build the Qwen community in Seoul with us
  • dcamp for hosting us at Seolleung
  • Cry Cheeseburger for the burgers
  • The volunteers who kept the night running: Bertijn, Brian, Nathaniel, Oscar, and Suhyun
  • And all 130 of you who came, asked unusually technical questions, and stayed around afterward to keep the conversations going

Watch the talks

Both full talks are online:

You can also find every published Dev Korea talk on our talks page.

See you at the next one.


Ready for your next move?
Visit Dev Korea to explore the latest job openings at dev-korea.com/jobs, or if you're hiring, post a job at dev-korea.com/post-a-job and connect with our growing international tech community.