Sunday, August 23, 2026

Off the Clock: Oh wie schön ist Panama

 



I'm a product-oriented software engineer who spends work hours thinking about products, users, and how to make things that matter. After hours, I work with my hands - fix things, build things, plant things, read things, and explore things. This is where those worlds meet. Sometimes a hobby teaches me something about my day job. Sometimes my product instincts change how I approach a craft. The lines cross more often than you'd expect. And this series of posts will be about that bi-directional influence. 


A Gem at the Library


A few weeks ago, I was at our local library in Berlin. Then a small children's book caught my attention. A bear and a tiger on the cover, looking happy. The title was "Oh, wie schön ist Panama". (how beautiful is Panama).

I read it many times. And every time I read it, I reflected more on some aspects in my life from their perspective.

Everything They Wanted


The story starts with the little bear and the little tiger. They live together in a small house by a river. They fish, they collect mushrooms, they have each other. Life is good. They have everything their heart desires. There is no problem to solve, no crisis to escape. They even have their own boat.

This reminded me of something I have seen many times in software teams. A system that works. Users are happily paying. The code is clean enough. Deployments are stable. And yet, someone reads a blog post or watches a conference talk, and suddenly the current setup does not feel enough anymore.

The Empty Box


One day, the bear finds a wooden crate in the river. The word "Panama" is written on it. The box is empty, but it smells of bananas. That is all. A word and a smell. From this single clue, the bear builds an entire dream in his head. Panama must be the land of their dreams. Everything there must be bigger, better, more beautiful, and smells like bananas.

In software, we do this all the time. We see a keyword on a landing page, a demo video, or a social media thread. Serverless. Microservices. AI-native. Blockchain. We do not investigate deeply. We do not ask if we actually need it. We just catch the smell of something exciting, and we build a whole dream around it.

The Departure


The bear goes home and tells the tiger about Panama. They talk late into the night. By morning, they have decided to leave. But there is one problem. They do not know which direction to go.

So they take the wooden crate and build a signpost. They place it at their door and point in one direction. They invented this direction themselves, but now they follow it with full conviction.

I have seen teams do exactly this. A working monolith gets torn apart into microservices that nobody needs. A simple CI pipeline gets replaced by a complex chain of tools. A stable hosting setup gets migrated to the next big cloud solution because everyone is doing it. The old system was not perfect, but it was working. The new dream was based on an empty box.

We build our own signposts. We read three articles, watch one keynote, and suddenly we know the direction. We point forward and start walking with confidence.

The Journey


Along the way, the bear and the tiger meet many animals. A fox, a crow, a hare, a hedgehog. Each one offers them something. A place to sleep, a story to hear, a new perspective. They discover many things they never knew before - like how a comfortable sofa feels. And they decide that one day they will have their own sofa too.

The journey is long and not always easy. They walk in the rain. They sleep in a barrel. They get hungry and look for food. But they keep going, because the dream pulls them forward.

In software, the migration journey is the same. You meet new tools, new communities, new ways of thinking. You learn about distributed systems, observability, infrastructure as code. You make connections. Some of these lessons are genuinely valuable. The problem is not what you learned. The problem is that you left home to learn it.

The Return


After a long journey, the bear and the tiger arrive at a place by a river. There is a small weather-beaten house. Big trees. A damaged bridge. A signpost on the ground which reads "Panama".

"This must be Panama!" they say. "It is even more beautiful than we imagined."

They do not recognize their own home. The wind and rain had changed it just enough. The overgrown trees made it look different. They see it with new eyes and fall in love with it.

Their boat is gone. It was destroyed during their absence. But they find the wood by the river and build a raft. Not the same as before, but something new they made themselves.

The Repair


The bear and the tiger do not sit and complain. They get to work. They fix the bridge. They repair the house. They clear the overgrown garden. They do this happily, because they believe they are building their dream.

This is what happens when teams finish a big rewrite or migration. They look at the new system and feel proud. It looks clean. It looks modern. But slowly, they would realize it solves the same problems the old one did. The business rules are similar. The challenges are familiar. They sought and built Panama and ended up at home. And sometimes, the old useful boat is gone. You build a new one from what you found along the way.

This is the part that changes everything. The story is not cynical. It does not say the journey was a waste. It says something warmer: sometimes you need to leave home to appreciate it. And when you come back, you bring the energy to fix what was always waiting for you.

In software, this is what real maintenance looks like. Not shameful debt reduction. Not grudging fixes. But caring for a system you now understand deeply because you tried everything else first.

The Sofa


A favorite detail in the story is the sofa. At the beginning, they did not have one. They only had wooden chairs. During their journey, they sat on someone else's sofa and loved it. When they arrived at their "Panama," they realized they wanted a comfortable one at home too.

They were delighted. But here is what they missed: they already had everything they needed. The sofa was just a nice addition to what was already complete.

The Lesson to Keep


Someone might ask if the bear and tiger could have just stayed home and saved themselves the long walk. The answer is no. Then they would never have met the fox, the crow, the hare, and the hedgehog. They would never have learned how their skills mattered and how they completed each other in the journey. They would never have seen their home from a new angle. And don't forget the Sofa!

The journey was not wasted. The return was not failure. Both were necessary.

And maybe that is the most honest thing I can say about software trends, framework migrations, and every empty box that smelled of bananas. We chase, we learn, we come back, and we fix what was always ours. Not because we were wrong. Because we needed the walk.


Friday, August 14, 2026

Off the Clock: Make a P.M. Decision



I'm a product-oriented software engineer who spends work hours thinking about products, users, and how to make things that matter. After hours, I work with my hands - fix things, build things, plant things, read things, and explore things. This is where those worlds meet. Sometimes a hobby teaches me something about my day job. Sometimes my product instincts change how I approach a craft. The lines cross more often than you'd expect. And this series of posts will be about that bi-directional influence. 

One of my favourite comedians is Kat Williams. And one of my fourth jokes he told tells a story about his son wanting a new Xbox. Instead, Kat convinced him to buy a used Nintendo console with two controllers and like 20 used games. His reasoning stuck with me for quite some time because it fits exactly how I think about product decisions in my engineering role.

He said multiple things in a matter of 2 minutes that changed how I see my work. Let me break down a few of them.

Daddy Can Do This All Day Every Day, No Problem

Kat said this about buying an Xbox for 199.99$. He had the money. He could do it. But he advised his son not to.

This sounded very close for engineers. We get asked to build things all the time. Use new frameworks. Add big fancy features. Complex architectures. My answer is often yes I can if that's the only choice. I can do it every day, all day. But that is not the question.

The real question is whether we should. A Product Manager mindset asks what outcome we are trying to reach. Is this the fastest path? Does this solve the actual problem? Or am I just showing off technical skills while wasting time?

I have said no to tasks I could easily complete. I once (or twice?) harmed my own bi-yearly performance review because I sounded "against the change" when it came to adding a new unnecessary fancy dependency to our stack. Not because I could not do them. Because I had a better idea for that time and effort. The best use of engineering power is not saying yes to everything. It is choosing the right thing to say yes to.

It Only Comes With 2 Demo Games and One Controller

Kat said the Xbox was a trap. Buy the base unit, then pay for every game. It comes with one controller. Daddy can't play with you. Your friends can't play with you. Buy another controller.

This happens in software too. A tool sounds free until you add the plugins. An API seems cheap until usage spikes. A microservice architecture starts simple but needs monitoring, logging, deployment pipelines, and five more services just to talk to each other.

These are hidden costs. Nobody shows them on day one. By the time you see them, you are already locked in. Migration takes months. Rewriting takes longer. Now you cannot go back.

A P.M. decision asks what comes in the box before you sign the contract. What are the real limits? What do I pay later that I ignore now?

20 Games That Other Kids Have Already Opened and Played With To Make Sure It Is Fun

Kat explained why the Nintendo was better. It came with like 20 games other kids had already tested. You knew they worked. You knew they were enjoyable.

Software has the same lesson. Some love shiny new tools. The latest library. The newest cloud service. The framework everyone is talking about. But who has really tested them yet in production? Nobody knows if they hold up under our pressure and our specific uses.

A P.M. decision looks for the games that already came with the system. Use patterns that work. Copy what others have learned. Solve the problem with code that has been proven before.

This does not mean staying old forever. It means being smart about when to risk. I would rather ship a boring feature that works than build something cool that crashes.

We Got Like 300 Dollars Leftover

If Kat's son chose the Nintendo over the Xbox, the saved money stayed in Kat's pocket. They could go across the country and have all the ice cream and nuggets they want! The point was not about saving cash. It was about having resources for other things.

In engineering we call this technical debt or opportunity cost. Every hour spent building unnecessary complexity is an hour taken from solving real user problems. Every server running unused features is budget gone from hiring someone new. Every bug in unfinished code slows down the next sprint.

Making a P.M. decision means looking at what you save when you choose simple over flashy. That saved time becomes freedom. Freedom to try real experiments. Freedom to pay down tech debt. Freedom to have fun building the next thing that actually matters.

Sometimes the smartest engineering move is doing less work.

The Bottom Line

Kat Williams taught his kid something bigger than video games. He taught him to look past the hype and the peer pressure. To value proof over promise. To keep resources for what matters.

Same applies to my work as an engineer. I can say yes to every request. I can build every feature. But a product mindset changes the question from can I to should I.

Make a P.M. decision. Choose the used game. Save the money. Keep the resources. Build what proves valuable over time.


Monday, July 13, 2026

Off the Clock: Prepping the Leather – Small Wins, Onboarding, and Lowering Barriers





I'm a product-oriented software engineer who spends work hours thinking about products, users, and how to make things that matter. After hours, I work with my hands - fix things, build things, plant things, read things, and explore things. This is where those worlds meet. Sometimes a hobby teaches me something about my day job. Sometimes my product instincts change how I approach a craft. The lines cross more often than you'd expect. And this series of posts will be about that bi-directional influence.



Some time ago, I had an idea: what if I sold a kit for making a leather wallet—not the full process, but just the final step?

I had already done the hard parts. I had measured, cut, punched holes, installed the push buttons, and glued the zipper. What remained was the stitching. The buyer would receive a nearly-complete wallet, two needles, some thread, and one skill to learn. One weekend, one satisfying project, one personal leather wallet.

The person who bought it told me it was a lovely weekend project. For them, the experience was about creating something without frustration. For me, it was about recognizing where the real barrier was—and removing everything except the one thing that makes the craft feel like craft.

What I didn't expect was how much this little experiment would influence how I think about my day job.

Where the barrier actually lives

When I started learning leathercraft myself, I ruined materials. Not because stitching is hard—it isn't, once you try it—but because getting to the stitching required a wall of tools, precise cuts, and mistakes that were expensive to make. The entry cost wasn't the skill. It was everything around the skill.

The same is true in software engineering. When a new engineer joins a team, the actual coding is rarely the bottleneck. The barrier is the setup, the domain knowledge, the unwritten rules, the context that lives in someone's head. By the time they write their first meaningful line of code, they've often already felt lost, slow, and unsure of themselves.

I've seen what happens when someone's first contribution is a small, successful one—fixing a bug, shipping a minor improvement. And I've seen what happens when their first task is unclear and goes nowhere. The first one builds momentum. The second one builds doubt.

The wallet kit taught me this in a different material: prep the hard parts, hand over the meaningful part, and let the learner feel competent early.

What product scoping and leather prepping have in common

There's a moment in leathercraft where you decide how far to prepare and where to stop. Cut too little and the user struggles. Prepare too much and you've removed the craft entirely—there's nothing left for them to own.

This is the same tension in product scoping.

When I've worked on launching a new product feature, we spent time in early research before writing a line of code. We looked at existing solutions, built small prototypes, and tested our assumptions. The goal wasn't to ship everything at once. It was to ship the minimum thing that would make a user feel: "Yes, this solves a real problem I have."

We cut features. Not because they weren't valuable, but because they weren't the first valuable thing. The wallet kit works because stitching is the first valuable thing—it's the moment where the pieces become a product. Everything before that is preparation. Everything after that is improvement.

In product, I've learned to ask the same question the wallet kit answers: what is the one action, the one moment of value, that this product delivers? Build toward that. Cut the rest for later.

The feedback loop I didn't plan for

Here's what surprised me: the product thinking I use at work started flowing back into my hobbies, and vice versa.

When I scope a product feature now, I think about the wallet kit. Not literally, but structurally: What am I asking the user to do before they reach value? Can I remove some of it? Am I handing them a pile of pieces and tools, or am I handing them a nearly-complete experience with one meaningful step to own?

When I design any other hobby project now, I think about user research. Who is this for? What do they already know? What will frustrate them, and is that frustration necessary—or is it just there because that's how I learned?

Neither side owns the lesson. They feed each other.

Small wins, bigger bets

The person who bought that kit wasn't trying to become a leatherworker. They wanted a weekend project that felt good. A small win.

But I've noticed—in engineering, in product, and in craft—that small wins are where bigger things begin. The new engineer who ships a fix on day three comes back on day five wanting something harder. The user who tries a feature and saves time asks what else it can do. The person who stitches a wallet starts looking at leather and wondering about more.

You don't get people to bigger commitments by making the first step harder. You get there by making the first step possible—and letting the satisfaction do the rest.

That's what the kit was really about. Not leather. Not wallets. Just the idea that lowering the barrier to a first win is one of the most generous things you can do—for a teammate, a user, or a stranger who wants to make something with their hands.

Tuesday, May 19, 2026

FinOps Beyond Cloud: Flagging Which LLM Path Runs

It’s 2026, and I pretend less that coding stops at merge. Plans still matter, but margin (revenue minus cost) matters too: if usage jumps, many users pile onto cheap tiers, or ARPU (average revenue per user) stays low, is your default technical path still affordable?

If you tilt product-minded, you often want costs and revenue to steer behavior—not just quarterly decks—and you ship guardrails (limits and alerts for spend and risk) plus metrics so typo fixes aren’t casually riding flagship AI tiers.

This post is a relaxed tour of one slice of that: treating cloud and model spend as something you design for, not something you discover on the invoice.

1. What a product-minded engineer optimizes beyond “merged”

Shipped is rarely the final step. The sharper test: multiply usage, push traffic through loss-leader tiers—does profitability still behave?

Checklist framing:

  1. Costs and revenue should sway day-to-day decisions.
  2. Use the cheapest good-enough path first; expensive models owe you justification.
  3. Add observability (logs/metrics showing what paths ran in prod) early so invoices don’t become the dashboard.

2. A profitability-aware seam—not just “flip the rollout switch”

Cost-aware feature flagging is more than switching release trains. (Feature flags toggle behavior remotely without reinstalling everywhere.) Core question shifts to:

For this payer plus this job step, does calling the pricey model earn its keep right now?

Two halves:

Inputs (FinOps-style facts): tooling that exposes spend (AWS Pricing CalculatorOpenCost), quotas (vendor-enforced caps), subscription revenue—not only beta_user=true.

Outputs: same outward job, staged execution—offline libraries; lighter Gemini Flash-class SKU; heavier Pro-class SKU when needed; delay/batch—not one greedy lane dialing the richest model every hop. (SKU or "Stock Keeping Unit" is vendor shorthand for a priced product bundle—the name on the invoice.)

Treat it like authentication (auth—checking who acts) guarding expensive work—but margin sits beside permission.


3. Models answer prompts—they don’t run budgets

A reader might ask whether frontier models magically cheap-route some of the simpler work.

Simply said: No. They optimize outputs within the SKU you bought; no accountants live there. You pick SKU tiers—they chase quality inside that sandbox. Toss a heavyweight tier at a petty job and it still replies, while metering charges requests, not vibes.

Operational decisions stay yours: routing (offline versus vendor lane), backoff (pause before retries when throttled), and hard-stop retry budgets. Models themselves won’t politely refund spend.


4. Total cost of ownership (TCO)—dashboards versus sloppy lanes

(TCO = total cost of ownership) includes quiet extras—not only billed tokens—like wrong tier ladders, retry storms, consolation “redo” generations after junk answers.

Log routing branch, approximate tokens (what vendors bill on) before stray heavyweight completions overshadow dashboard cost. Procurement warning ahead: illustrative magnitudes only—junk routing has landed teams roughly two to four orders of magnitude (~100×–10,000×) hotter than guarded paths.

Prefer modest telemetry before invoices swell.


5. From toy splitter to nearer-production routers

Toy setup uses HTTP POST /v1/check with JSON { text, task } and tasks spellcheck or summarize.

Environment COST_OPTIMIZATION_ENABLED when on triggers Typo.js (open-source English spellchecking) plus an extractive summary (reuses existing sentences—invents nothing).

When offGemini 2.5 Flash (Google positions this as its lighter SKU string gemini-2.5-flash) owns both passes.

Closer to prod, swap the simple “if-this-then-that” rules for nuanced routing plus subscription and billing knobs—still feeding FinOps signals.


6. Sampling experiment—methods first; numbers follow

This section lays out (1) what mechanically reran, (2) tables in later subsections, (3) how extrapolation pencils out—lab notebook plus spreadsheet honesty, not magic.

6.1 What reran—and what got skipped deliberately

Bench here means scripted automation issuing the same deterministic requests.

Ran: fabricated multi-section write-up (~290 words) with sequential jobs (spellcheck then summarize). Express route POST /v1/check. Flag COST_OPTIMIZATION_ENABLED jointly toggles offline Typo stack versus Gemini 2.5 Flash via @google/genai. Then a script prints timings plus playful dollar placeholders—not accounting truth.

Skipped: proportional live traffic mixtures, randomized A/B tests (A/B: compare flows across user subsets)—this isolate covers mechanism, not population behavior.

Driving question stays:

Holding document shape steady and accepting cheaper summaries, how many vendor completions vanish?

Extrapolation (spoken plainly):

Rough monthly dollars skipped ≈ (cloud completions you avoided) × (realistic dollars per completion).

Assume each workflow run equals the spell+summarize pair (two HTTP POSTs)—linear scaling versus monthly runs. Lock those two POSTs/run, multiply by how many drafts, audits, or refactors fire monthly.

6.2 Head-to-head results (the sample)

MetricOptimization ONOptimization OFF
Requests22
Offline routes20
Cloud (Gemini)02
Avg latency (ms)~2378~4788
Placeholder session total$0.000000$0.000230

The session total is only the demo knobs COST_CLOUD_SPELLCHECK + COST_CLOUD_SUMMARY in server.js—fine for a workshop, not a bill.

Plain-language takeaway for this slice: optimization on skipped every Gemini invocation that off made for these two tasks (0 vs 2 calls → 100% avoidance for this pair only). Tradeoff: extractive bullets and a full-file dictionary pass are cheap, but the spell step can land around a few seconds on a long draft (~4.7s on the first spellcheck request in my run), while the offline summary stayed single-digit ms.

6.3 Projecting to volume (drafts, edits, audits, refactors)

Illustrative token sketch—same Flash-style guesses baked into the benchmark (~2200 / 2000 in/out for the long corrective pass, ~2200 / 450 for exec summary, $0.15/M in and $0.60/M out; verify Gemini pricing). Think of a run as one time you execute spellcheck + summarize on a body—e.g. a new draft, a heavy edit, an audit pass that re-queues the pair, or a refactor that rewrites a section and re-runs the tooling.

ItemOptimization ON (this demo)Optimization OFF (all Gemini for these two tasks)
Cloud calls / month @ 5k runs (2 POSTs per run)010,000
Order-of-magnitude API $ / month~$0~$10.65
Avoided vs all-cloud at that volume (token model only)~$10.65

Scale the numerator: at 50k runs/month—stacking draft cycles, revision rounds, audit and compliance re-checks, refactors, anything that replays this two-step cloud path—the same all-cloud token math lands near ~$106/month for that slice alone—before you add any new AI feature. The cost-aware pattern is convincing here because the growth lever is obvious: more passes through the pipeline ⇒ more invocations ⇒ the same percentage of avoided calls buys proportionally more dollars as activity grows.

6.4 Where product growth quietly multiplies cost (features, not just users)

Users rarely stop at “spellcheck + summary.” Roadmaps add adjacent model tasks: e.g. AI glossary (“explain this acronym for a non-security exec”), keyword callouts next to the summary, a risk-language nudgeemail-ready rewrite, or second-pass tightening. If each is implemented as another always-on Gemini call, you get multiplication: two cloud tasks become three, four, five—each time someone runs the flow—while routing stays an afterthought.

A cost-aware seam doesn’t mean shipping worse product; it means deciding per task (plan tier, cache, template, small model, batch, human review) instead of defaulting “new AI affordance ⇒ new flagship invocation.” The bench only models two tasks, but the projection mental model extends: every new task is a coefficient on monthly variable spend unless you fold it into the same router.

6.5 Why this is still a convincing case for the approach

  1. The sample isolates mechanism: you can see exactly which HTTP paths hit the model when the flag flips—no mystery meat in “optimization.”
  2. Volume makes small per-call numbers real: ~$0.00107/POST on the all-cloud token estimate at 10k calls/month is easy to shrug off until draft/edit/audit/refactor volume (and features) push you to 100k+ calls.
  3. Feature creep is the hidden multiplier: routing discipline is how you ship more AI-shaped surface area without linear-to-cloud spend on every new button.


Please note: The playground code for server.js is at the bottom of this post. The post treats that run as a sample you can scale with your own monthly pipeline volume—how often teams hit spellcheck + summarize across drafts, edits, audits, refactors, and so on—and your task list. Dollar figures mix placeholder session costs and illustrative token mathabsolute savings scale with users, calls per run, model tier, and how many new AI features stay cloud-default—plug in real metering before treating any number as financial guidance.


require('dotenv').config();
const express = require('express');
const { GoogleGenAI, ApiError } = require('@google/genai');
const Typo = require('typo-js');

const app = express();
app.use(express.json());

const dictionary = new Typo('en_US');

const genai = new GoogleGenAI({
    apiKey: process.env.GEMINI_API_KEY,
});

/** Placeholder $ per cloud call (tune for FinOps demos; real bills use metering). */
const COST_CLOUD_SPELLCHECK = Number(process.env.COST_CLOUD_SPELLCHECK) || 0.00005;
const COST_CLOUD_SUMMARY = Number(process.env.COST_CLOUD_SUMMARY) || 0.00018;

/** Correct spelling per word; preserves whitespace and punctuation (en_US). */
function correctWithTypo(text) {
    return text.replace(/\b[\w']+\b/g, (word) => {
        if (dictionary.check(word)) return word;
        const suggestion = dictionary.suggest(word)[0];
        return suggestion || word;
    });
}

/** Cheap path: lead sentences + pseudo-bullets (no API). */
function extractiveExecutiveSummary(text) {
    const t = text.trim().replace(/\s+/g, ' ');
    const sentences = t.split(/(?<=[.!?])\s+/).filter((s) => s.length > 15);
    const head = sentences.slice(0, 4).join(' ');
    const base =
        head.length >= 200 ? head : t.slice(0, Math.min(1200, t.length)) + (t.length > 1200 ? '…' : '');
    const lines = base
        .split(/(?<=[.!?])\s+/)
        .filter(Boolean)
        .slice(0, 5)
        .map((s) => `- ${s.trim()}`);
    return lines.join('\n');
}

let sessionTotalCost = 0;

function isTruthyEnv(name) {
    const v = process.env[name];
    if (v == null) return false;
    return /^(1|true|yes)$/i.test(String(v).trim());
}

/** Pull a readable message out of SDK errors (often `message` is stringified JSON). */
function geminiErrorDetail(err) {
    const msg = err && typeof err.message === 'string' ? err.message : String(err);
    try {
        const parsed = JSON.parse(msg);
        const inner = parsed && parsed.error ? parsed.error : parsed;
        if (inner && typeof inner.message === 'string') {
            return { summary: inner.message, code: inner.code, status: inner.status };
        }
    } catch (_) {
        /* use raw */
    }
    return { summary: msg };
}

app.post('/v1/check', async (req, res) => {
    const { text, task } = req.body ?? {};
    if (typeof text !== 'string' || !text.trim()) {
        return res.status(400).json({ error: 'Body must include non-empty string `text`.' });
    }
    if (task !== 'spellcheck' && task !== 'summarize') {
        return res.status(400).json({
            error: 'Body must include `task`: "spellcheck" | "summarize".',
        });
    }

    const isOptimizationOn = isTruthyEnv('COST_OPTIMIZATION_ENABLED');
    const words = text.trim().split(/\s+/);

    let output = '';
    let engine = '';
    let cost = 0;

    try {
        if (task === 'spellcheck') {
            if (isOptimizationOn) {
                if (words.length <= 5) {
                    output = correctWithTypo(text);
                    engine = 'Offline (Typo.js)';
                } else {
                    output = correctWithTypo(text);
                    engine = 'Offline (Typo.js full document)';
                }
                cost = 0;
            } else {
                const model = process.env.GEMINI_MODEL || 'gemini-2.5-flash';
                const response = await genai.models.generateContent({
                    model,
                    contents: [
                        'You correct spelling and obvious typos only. Preserve structure, headings, and meaning. Reply with the full corrected text only—no preamble or quotes.',
                        `Document:\n${text}`,
                    ].join('\n\n'),
                });
                const raw = response.text;
                if (raw == null || !String(raw).trim()) {
                    throw new Error('Empty response from model');
                }
                output = String(raw).trim();
                engine = `Cloud spellcheck (${model})`;
                cost = COST_CLOUD_SPELLCHECK;
            }
        } else {
            // summarize → executive summary
            if (isOptimizationOn) {
                output = extractiveExecutiveSummary(text);
                engine = 'Offline (extractive executive summary)';
                cost = 0;
            } else {
                const model = process.env.GEMINI_MODEL || 'gemini-2.5-flash';
                const response = await genai.models.generateContent({
                    model,
                    contents: [
                        'Condense the report into an executive summary for leadership: 3–5 bullet points, plain text, each line starting with "- ". Be factual; do not invent risks or metrics.',
                        `Report:\n${text}`,
                    ].join('\n\n'),
                });
                const raw = response.text;
                if (raw == null || !String(raw).trim()) {
                    throw new Error('Empty response from model');
                }
                output = String(raw).trim();
                engine = `Cloud summary (${model})`;
                cost = COST_CLOUD_SUMMARY;
            }
        }

        sessionTotalCost += cost;
        console.log(`[${new Date().toISOString()}] task=${task} Engine: ${engine} | Cost: $${cost}`);

        res.json({
            task,
            output,
            engine,
            stats: { sessionTotal: sessionTotalCost.toFixed(6) },
        });
    } catch (error) {
        if (error instanceof ApiError) {
            const { summary, code, status: bodyStatus } = geminiErrorDetail(error);
            const httpStatus =
                typeof error.status === 'number' && error.status >= 400 && error.status <= 599
                    ? error.status
                    : 502;
            console.error('[Gemini ApiError]', httpStatus, summary);
            const label =
                httpStatus === 429
                    ? 'Gemini quota or rate limit (check plan / AI Studio quotas)'
                    : 'Gemini API error';
            return res.status(httpStatus).json({
                error: label,
                details: summary,
                geminiCode: code,
                geminiStatus: bodyStatus,
            });
        }
        console.error('[Report pipeline]', error);
        res.status(500).json({ error: 'System Error', details: error.message });
    }
});

const PORT = Number(process.env.PORT) || 3000;

const server = app.listen(PORT);
server.once('listening', () => {
    console.log(`Report spellcheck + executive summary API on port ${PORT}`);
});
server.once('error', (err) => {
    console.error('Server failed to start:', err.code === 'EADDRINUSE' ? `port ${PORT} is already in use` : err.message);
    process.exit(1);
});



Monday, April 20, 2026

Local AI on an M1 Pro: From "Post-Apocalyptic" Slowness to a Functional Reality

We’ve all seen headlines like this: "The end of paid coding assistants!" or "Run your own private AI locally for free!" As someone who values privacy and hates the dependency on monthly subscriptions, I decided to see if my trusty MacBook Pro (M1 Pro, 32GB RAM) could become my trustworthy coding workstation.

My journey started with a massive failure, moved through a "thinking" loop, and finally landed on a configuration that actually works. Here is how I turned my Mac into a functional local coding station.

(Yeah... I’ve reached the stage where I won’t bother with perfect screenshots. If I need to show the output of my fun tinkering on an isolated machine, I’ll just take a blurry photo of the screen with my old phone. It’s much more relaxed that way.)



Phase 1: The "Heavyweight" Disaster with "Qwen 3.5 Coder Next"

I started ambitious. Installing Ollama was the simplest part. Then I pulled Qwen 3.5 Coder Next. I thought 32GB of RAM would be enough. I was wrong.

The Experience:

  • Initial 'Hi': 25 seconds.

  • Coding Task: 7+ minutes to generate just some initial instructions and start writing the variables section for an Arduino script.

  • The Culprit: My logs showed the model needed 51.3 GiB of memory. Since I only have 32GB, Ollama had to shove 26GB of the "brain" onto my CPU.


The long wait after a simple "Hi".

 

The painful 7+ minutes of waiting while watching the slow code generation and seeing my computer's memory heavily consumed.



Phase 2: The "Thinking" Trap with "Qwen 3.5 9B"

I pivoted to a smaller model: Qwen 3.5 9B. On paper, this should have been lightning fast. However, I ran into a new hurdle: Reasoning Loops.

Even with the smaller 9B model, the "Reasoning/Thinking" phase was taking forever—sometimes up to 7 minutes of "thinking" without a single line of code being written. At one point, it even got caught in a logical loop, and I had to restart the process.

A lot of thinking for 7+ minutes and no action but much less memory consumption.



Phase 3: The Breakthrough (The "Nothink" Secret)

The real "Aha!" moment came when I realized I didn't need the model to spend several minutes pondering the meaning of life for a simple Arduino script. I just needed the code.

I used a simple command to bypass the heavy reasoning phase: >>> /set nothink

The difference was huge:

  • Total Response Time: 2 minutes and 28 seconds for a complete, complex answer.

  • Content Quality: It wasn't just code. It gave me prerequisites, circuit wiring, full Arduino code, security tips, and even improvement ideas.

  • Memory Efficiency: The logs show this model is a perfect fit for the M1 Pro. It only used about 9.1 GiB of total memory, meaning 100% of the model layers (33/33) stayed on the GPU (Metal).

This time, with 'nothink', it was spitting out the answer much faster.



Technical Insights from the Logs

If you are troubleshooting your own local setup, here is what I learned from the Ollama server logs:

  1. Check your Offloading: In my successful 9B run, the logs said: offloaded 33/33 layers to GPU. This is the ideal case. If that number isn't 100%, your performance will be affected.

  2. Flash Attention is King: The logs confirmed enabling flash attention. This helps the Mac handle long conversations without slowing down.

  3. The "Unified" Advantage: My M1 Pro was able to allocate a recommendedMaxWorkingSetSize of ~26GB. By using the 9B model (which only needs ~9GB), I left plenty of room for my system to breathe.


The Verdict: Is it a "Free" Coding Assistant?

Is a "free" local coding assistant possible? Yes—but resources matters. If you try to run massive models on a 32GB Mac, you'll feel like you're back in the era of dial-up.

If this continues to feel this good, I’m considering the ultimate "pro" move: Using this Mac as a headless AI server. I can connect it to a secure network and put it away on a shelf and connect to its "brain" from my other computers for chatting or from within VS Code (with the suitable plugin's of course) or even my phone. My main coding machine stays cool and quiet, while the M1 Pro does all the heavy lifting in the background.

I hope that my little weekend experiment has helped you in any way with insights or inspiration to set on your own journey of finding your own AI independence.


Tuesday, April 7, 2026

The Rise of the Product Engineer: Title Trend or New Reality?

 

This career topic has been on my mind for a while, and I've been trying to collect more information about it for my own career's sake. And now I think I have an idea clear enough to be shared. I hope it benefits someone out there.

The "Product Engineer" title has exploded in popularity this year. This shift is happening for three main reasons:

  • AI-Assisted Coding: Since AI can now handle basic coding tasks, companies need engineers who can move "above" the code to control the requirements (inputs) and the architecture (outputs).

  • Flatter Teams: Tech companies are removing middle layers, requiring engineers to be more independent.

  • Faster Delivery: To move quickly, the line between "thinking about the product" and "writing the code" must disappear.

However, after talking to managers, recruiters, and "Product Engineers", I realized that not everyone defines this role the same way. Here are the three main types of "Product Engineers" I have observed:

1. The "Label Switch" (Product in Name Only)

In these companies, the title is just a marketing trick. They swapped the word "Software" for "Product," but nothing else changed.

  • The Reality: Whether you are a junior or a senior, your job is the same as a traditional "heads-down" coder.

  • The Hiring Process: The interview is 100% technical. They don't assess your business knowledge or how you think about users.

  • The Day-to-Day: You receive a ticket, you code it, and you move on. The "product" part is just a fancy new sticker on your LinkedIn profile. You might negotiate the requirements or the sequence of shipping things with your PM, but that's the same thing as the past decades in any small to medium startup. Nothing new.

2. The Product-Minded Engineer

This is a more mature approach, often seen in tech companies with a transparent and flexible management style. Here, the engineer is a partner to the Product Manager (PM).

  • The Reality: Mid-level and senior engineers are expected to help improve requirements, not just follow them. There is a heavy focus on customer value over "tech talk."

  • The Hiring Process: Interviews include a specific section to discuss how you’ve solved user problems, how you collaborate with designers, and how you interact with the PMs in earlier stages.

  • The Day-to-Day: About 10% of your time is spent on product strategy. You are a pragmatist who knows when to choose a "good enough" technical solution to help the user faster. However, the PM still holds the final accountability for the roadmap.

3. The "Part-Time PM" Engineer

This is the most intense version of the role and the unicorn of that job title. These companies need someone who can lead a project from a blank page to a finished product.

  • The Reality: You are essentially a Product Manager who also writes code. You are responsible for the "Why" and the "How."

  • The Hiring Process: Be prepared for deep questions about product frameworks, data analysis, and user research. They want to see if you can lead a squad of engineers.

  • The Day-to-Day: You participate in ideation, talk to stakeholders, and conduct user interviews. You shape the work for the rest of the team and ensure the technical output matches the business goals perfectly. Expect extra accountabilities with this version.


Conclusion

The software industry is moving away from "coding as a service" toward "problem-solving as a service." Depending on the company, a Product Engineer can be a simple developer or a business leader. If you are looking for this role, make sure to ask during the interview: "How much influence do I actually have over the 'Why' of the product? And am I actually accountable for any decision made?"