<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://gabrielkasmi.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://gabrielkasmi.github.io/" rel="alternate" type="text/html" /><updated>2026-09-15T22:16:32+00:00</updated><id>https://gabrielkasmi.github.io/feed.xml</id><title type="html">Gabriel Kasmi</title><subtitle>ML scientist working on explainable AI, computer vision, and remote sensing for energy systems.</subtitle><author><name>Gabriel Kasmi</name></author><entry><title type="html">Navier–Stokes, Scaling Laws, and Peak Tokens</title><link href="https://gabrielkasmi.github.io/blog/2026/09/15/scaling-laws-navier-stokes/" rel="alternate" type="text/html" title="Navier–Stokes, Scaling Laws, and Peak Tokens" /><published>2026-09-15T00:00:00+00:00</published><updated>2026-09-15T00:00:00+00:00</updated><id>https://gabrielkasmi.github.io/blog/2026/09/15/scaling-laws-navier-stokes</id><content type="html" xml:base="https://gabrielkasmi.github.io/blog/2026/09/15/scaling-laws-navier-stokes/"><![CDATA[<p>
  Much has already been said about the OpenAI Navier–Stokes affair — about the credibility of the method itself (does the model actually reason, or is this brute force dressed up as insight), and about the credit dispute with the mathematicians who were working the same problem. One detail, though, is worth dwelling on longer than the rest of the coverage has: OpenAI's own admission that <em>"while unlikely, [it] cannot rule out that de-identified data derived from their usage of our products helped improve our models."</em>
</p>

<p>
  Behind that sentence sits a fairly simple fact: any AI model needs good data to keep improving. The public conversation about AI's limits keeps circling around an energy wall, or a valuation bubble. But what if the real binding constraint is something narrower and more specific — tokens — and what if, for this particular resource, the race to the frontier is set up to cannibalize itself?
</p>

<h3>1. Data is all you need</h3>

<p>
  At first order, progress in these models has depended on scaling two things together: compute and training tokens. Very roughly speaking, for each additional parameter, 20 additional tokens are needed.
</p>

<p>
  This arithmetic quickly exhausted the internet. Current models are already trained on the whole internet. It's worth noting, in passing, that the <em>other</em> commonly discussed constraint — energy and compute infrastructure — is the one the industry actually knows how to throw engineering at, however extreme: Google and SpaceX are jointly developing Project Suncatcher, orbital data centers running on continuous solar power, targeting deployment as early as 2027. There is, notably, no equivalent engineering fix for a shortage of genuinely new ideas. You cannot launch a satellite to go fetch a thought that doesn't exist yet.
</p>

<h3>2. The law of receding tokens</h3>

<p>
  With the open internet drained, the hunt has moved to other reservoirs: open data commons that were never built to withstand this kind of load. Wikimedia has documented a 50% surge in bandwidth demand on Wikimedia Commons driven by AI crawlers; The Register names Meta and OpenAI specifically as the worst offenders. The same pressure is reported anecdotally around other open infrastructure — OpenStreetMap and Zenodo among them — maintainers describing a need to defend against "excessive automation" alongside legitimate use.
</p>

<p>
  But the open commons aren't the only reservoir left to mine. The Navier–Stokes episode reveals the next one: users' own prompts. If interactions with a model can measurably improve it, then prompts are training data too — original, unpublished, often more valuable per token than anything left on the open web, precisely because they capture reasoning that has never been written down anywhere else. The question this opens isn't legal so much as reputational: once people understand this is happening, how will it be received?
</p>

<h3>3. The clock is ticking</h3>

<p>
  If the pool of good prompts is the last frontier, it's also a strikingly fragile one, for at least two reasons.
</p>

<p>
  First, an eviction effect: people with genuinely original ideas may simply stop feeding them to a tool, once the Navier–Stokes precedent is out there. This isn't a hypothetical reaction — the mechanism has already been measured elsewhere. After the Snowden revelations, Jon Penney's study of Wikipedia traffic found a measurable drop in visits to sensitive topics, driven purely by the awareness of being watched, without any actual enforcement action ever taking place. There's no reason to expect researchers and domain experts to behave differently once they suspect their brainstorming sessions might be harvested.
</p>

<p>
  Second, a structural response: the growth of self-hosted, open-weight models will likely absorb a growing share of exactly this kind of high-value brainstorming, run locally and never reaching Anthropic's or OpenAI's servers at all.
</p>

<figure class="blog-figure">
  <img src="/blog/img/scaling-laws-navier-stokes/goya-saturn.jpg" alt="Francisco de Goya, Saturn Devouring His Son" style="display:block; max-width:55%; height:auto; margin:0 auto;" />
  <figcaption>Francisco de Goya, <em>Saturno devorando a su hijo</em> (c. 1820). Public domain (Wikimedia Commons).</figcaption>
</figure>

<p>
  OpenAI's celestial naming scheme is well chosen, incidentally — Terra, Luna, Sol. A fitting next name for GPT-6.1 might be Saturn: the Roman god who devoured his own children. More fundamentally, though: it's worth asking whether the deposit of usable data needed to train the next generation of these models might run dry before the data centers meant to train them have even finished being built.
</p>

<h3>Sources</h3>

<ul>
  <li><a href="https://openai.com/index/navier-stokes-solution/">OpenAI — On the Navier–Stokes Millennium Prize Problem</a></li>
  <li><a href="https://arxiv.org/html/2406.12907v1">Reconciling Kaplan and Chinchilla Scaling Laws (arXiv)</a></li>
  <li><a href="https://www.pbs.org/newshour/economy/ai-gold-rush-for-chatbot-training-data-could-run-out-of-human-written-text-as-early-as-2026">PBS NewsHour — AI gold rush for chatbot training data could run out of human-written text as early as 2026</a></li>
  <li><a href="https://diff.wikimedia.org/2025/04/01/how-crawlers-impact-the-operations-of-the-wikimedia-projects/">Wikimedia Diff — How crawlers impact the operations of the Wikimedia projects</a></li>
  <li><a href="https://techcrunch.com/2025/04/02/ai-crawlers-cause-wikimedia-commons-bandwidth-demands-to-surge-50">TechCrunch — AI crawlers cause Wikimedia Commons bandwidth demands to surge 50%</a></li>
  <li><a href="https://www.theregister.com/2025/08/21/ai_crawler_traffic/">The Register — AI crawlers and fetchers are blowing up websites, with Meta and OpenAI the worst offenders</a></li>
  <li><a href="https://pretalx.com/sotm2026/speaker/YGZZPU/">State of the Map 2026 — "Running OpenStreetMap.org in the Age of AI" (Grant Slater)</a></li>
  <li><a href="https://theconversation.com/data-centres-in-space-will-2027-really-be-the-year-ai-goes-to-orbit-271018">The Conversation — Data centres in space: will 2027 really be the year AI goes to orbit?</a></li>
  <li><a href="https://www.datacenterdynamics.com/en/news/project-suncatcher-google-to-launch-tpus-into-orbit-with-planet-labs-envisions-1km-arrays-of-81-satellite-compute-clusters/">Data Center Dynamics — Project Suncatcher</a></li>
  <li><a href="https://theintercept.com/2016/04/28/new-study-shows-mass-surveillance-breeds-meekness-fear-and-self-censorship/">The Intercept — Mass surveillance breeds meekness, fear, and self-censorship, new study shows (Jon Penney)</a></li>
</ul>]]></content><author><name>Gabriel Kasmi</name></author><category term="AI" /><category term="Scaling Laws" /><category term="Open Data" /><summary type="html"><![CDATA[OpenAI's Navier–Stokes result quietly admits that user prompts may be training data too. A look at why tokens — not energy or money — could be AI's binding constraint, and why that last reservoir is set up to run dry.]]></summary></entry><entry><title type="html">Hidden in Plain Sight</title><link href="https://gabrielkasmi.github.io/blog/2026/09/04/hidden-in-plain-sight/" rel="alternate" type="text/html" title="Hidden in Plain Sight" /><published>2026-09-04T00:00:00+00:00</published><updated>2026-09-04T00:00:00+00:00</updated><id>https://gabrielkasmi.github.io/blog/2026/09/04/hidden-in-plain-sight</id><content type="html" xml:base="https://gabrielkasmi.github.io/blog/2026/09/04/hidden-in-plain-sight/"><![CDATA[<p>
  Like most people, I had used OpenStreetMap for years without ever really understanding the idea underneath it. It was just "the open Google Maps", with a lot of large systems already referenced in Europe.
</p>

<p>
  It took a weekend at State of the Map to realize that OSM rests on two core principles, far more interesting than the map itself.
</p>

<p>
  The first: <strong>you only map what can be seen.</strong> OSM records what is observable to anyone standing in a public place: a road, a roundabout, a building's footprint. The second: <strong>the map cannot be appropriated.</strong> It's a commons. No company, no state, no individual can enclose it. Everyone contributes, everyone benefits, and the thing itself belongs to no one.
</p>

<p>
  Two modest-sounding rules. I've come to think they're the answer to a discomfort that keeps coming up in my own work, and in a debate I watched at the conference.
</p>

<h3>Panoramax, or: we put photos on the internet</h3>

<p>
  One of the flagship projects in the room was <strong>Panoramax</strong>, an open, community-run alternative to Street View. Volunteers capture street-level imagery and publish it openly for anyone to use.
</p>

<p>
  And immediately, the uncomfortable question: we are photographing streets and putting them online. A parked car that moves at a certain hour tells you someone has left home. Multiply that across a city, point modern AI at it, and you extract a lot that people never meant to share. So, do we <em>want</em> to enable that?
</p>

<p>
  It's a fair question. But notice where it bites. The car was on the street, in plain view, the whole time. The imagery didn't create the information; it made it <em>easier to collect at scale</em>. The line that matters isn't between <strong>taking the photograph</strong> and <strong>not taking it</strong> — it's between taking it and <strong>misusing it</strong>.
</p>

<h3>"You're going to look inside people's homes"</h3>

<p>
  I recognize that discomfort because it followed me all through my PhD. Every time I described mapping rooftop solar from aerial imagery, someone would say some version of: <em>aren't you going to look inside people's homes?</em>
</p>

<p>
  It's the same worry as Panoramax, wearing different clothes. A solar panel on a roof is visible from the sky, and it always has been. What changed is that a model can now find all of them, everywhere, cheaply. The instinct, in both cases, is the same: this feels intrusive, so maybe we should keep the data closed.
</p>

<h3>The counter-intuitive part</h3>

<p>
  Here's the thing about closing a dataset of something already visible: <strong>it doesn't hide the phenomenon. It only decides who gets to see it.</strong>
</p>

<p>
  The rooftop is still there. The street is still filmed by every passing dashcam. Locking away the <em>map</em> doesn't remove the <em>capability</em> — it concentrates it in the hands of whoever can afford to rebuild it privately, while denying it to the researchers, local authorities and citizens who would use it in the open.
</p>

<p>
  Cryptographers have a name for the mistake this avoids. <strong>Kerckhoffs's principle</strong> says a system should stay secure even if everything about it <em>except the key</em> is public. Security that depends on nobody knowing how the system works isn't security — it's a fragility waiting to be found.
</p>

<p>
  Translate that to infrastructure data and it reads almost the same:
</p>

<blockquote>
  <p><em>A system you have to keep secret to keep safe isn't safe, it's fragile. What's robust is what everyone can see, check, and correct.</em></p>
</blockquote>

<p>
  This is where OSM's two principles click into place. <strong>Map only what's visible</strong>, and you're not building surveillance — you're describing what is already public. <strong>Make it inappropriable</strong>, and no single bad actor gains an asymmetric edge, because everyone holds the same map, including the people who would spot and fix its errors. You regulate the <em>misuse</em>, not the <em>looking</em>.
</p>

<figure class="blog-figure">
  <img-comparison-slider style="width:100%; --divider-width:2px;">
    <img slot="first" src="/blog/img/hidden-in-plain-sight/before.webp" alt="Aerial view over Montélimar — before detection" width="1600" height="890" style="width:100%; display:block;" />
    <img slot="second" src="/blog/img/hidden-in-plain-sight/after.webp" alt="The same view with detected rooftop PV installations" width="1600" height="890" style="width:100%; display:block;" />
  </img-comparison-slider>
  <figcaption>It was there all along — drag the slider to reveal the detected rooftop PV. Detections over Montélimar, from <a href="https://deeppvmapper.fr">deeppvmapper.fr</a>.</figcaption>
</figure>

<h3>The case for open PV mapping</h3>

<p>
  This isn't abstract for me. Rooftop solar is a near-perfect example of something <em>hidden in plain sight</em>: visible from space, yet largely missing from the official registries meant to track it. In several countries the gap between what's installed and what's recorded runs into the tens of percent.
</p>

<p>
  You can close that gap two ways. You can buy exclusive high-resolution imagery, run it privately, and publish a number nobody else can reproduce. You get a figure standing on sand, trustworthy only as far as you trust the people who made it. Or you can do it in the open.
</p>

<p>
  That's the bet behind <strong><a href="https://open-energy-transition.github.io/earthpv/">EarthPV</a></strong> and <strong><a href="https://deeppvmapper.fr/">DeepPVMapper</a></strong>: open imagery, open models, open data, and a crowdsourcing loop where anyone can verify a detection and feed the correction back in. The map isn't weaker for being public. Every extra pair of eyes that checks a panel, every community that corrects a boundary, makes it better. It's a public good in the literal sense: it gains value as more people use it.
</p>

<p>
  I'm not trying to settle the grand "should data be open or closed" debate. My point is narrower, and I think harder to argue with: for infrastructure <em>already visible from the sky or the street</em>, the real choice was never "seen or hidden." It's "seen by a well-resourced few, or seen — and correctable — by everyone." Openness isn't the risky option here. It's the robust one.
</p>]]></content><author><name>Gabriel Kasmi</name></author><category term="Open Data" /><category term="OpenStreetMap" /><category term="Photovoltaics" /><summary type="html"><![CDATA[Something visible from the sky or the street isn't made safe by hiding the dataset — that only decides who gets to see it. Why openness is the robust choice for mapping infrastructure like rooftop solar.]]></summary></entry><entry><title type="html">The Blind Men and the Solar Panel</title><link href="https://gabrielkasmi.github.io/blog/2026/08/15/three-maps-three-definitions/" rel="alternate" type="text/html" title="The Blind Men and the Solar Panel" /><published>2026-08-15T00:00:00+00:00</published><updated>2026-08-15T00:00:00+00:00</updated><id>https://gabrielkasmi.github.io/blog/2026/08/15/three-maps-three-definitions</id><content type="html" xml:base="https://gabrielkasmi.github.io/blog/2026/08/15/three-maps-three-definitions/"><![CDATA[<p>
  MapYourGrid recently opened a <a href="https://wiki.openstreetmap.org/wiki/Proposal:Power_generation_storage">proposal to overhaul how power generation and storage are tagged in OpenStreetMap</a>. It's the first real revision since 2013, when the <a href="https://wiki.openstreetmap.org/wiki/Proposal:Power_generation_refinement">current scheme</a> — was put in place. On the surface it's a fairly dry tagging debate: new keys for <code>source</code>, <code>method</code>, <code>technology</code>, the ability to chain generators together, explicit roles like <code>main</code>, <code>backup</code>, <code>auxiliary</code>, <code>standalone</code>, a proper way to describe storage.
</p>

<p>
  This debate is not only a question for cartographers. Every convention defines what a map is capable of showing — and, by construction, about what it can't. A registry, a satellite pipeline, and a crowdsourced map are not just three ways of drawing the same picture; they're three different answers to the question of where our blind spots sit when we try to monitor energy infrastructure. Get the object definition wrong, or leave it implicit, and the blind spot doesn't disappear — it just becomes invisible to whoever is relying on the map.
</p>

<p>
  Photovoltaics is a good illustration, because it's far less simple to define than it looks. And the fact that OSM's own community is still renegotiating this, more than a decade after the last attempt, is a signal in itself. The tension we run into when comparing grid-connection data against remote sensing isn't a quirk of our particular pipeline — it's a modeling problem for the object "PV installation" itself, one that resurfaces even inside an open, collaboratively negotiated system built specifically to standardize this kind of thing.
</p>

<p>
  Three groups routinely try to describe this object: grid operators, remote sensing pipelines, and open mapping communities. Each has a real, working definition. Each is right, given what it's built to see. None of them is describing the whole thing.
</p>

<p>
  It turns out PV is the elephant of the old parable — the blind men and the elephant, updated for the energy transition. A grid operator has a hand on the tail. A satellite pipeline is feeling up the ear. A mapper is patting the trunk, in the dark, one contribution at a time. Each comes back with a description that's internally consistent, confidently reported, and completely partial. Nobody is lying. Nobody has the whole animal.
</p>

<figure class="blog-figure">
  <img src="/blog/img/three-maps-three-definitions/elephants.webp" alt="The blind men and the elephant, reimagined with a grid operator, a satellite, and a mapper each examining a different part of the same PV installation" />
</figure>

<h3>1. A blurry ground truth</h3>

<p>
  Before comparing <em>representations</em> of PV installations, it's worth asking whether the <em>thing being represented</em> holds still long enough to be defined.
</p>

<p>
  Start with the easy case that turns out not to be so easy. <strong>Roof-mounted versus ground-mounted</strong> looks like a clean line — until you get to <strong>solar carports</strong> (ombrières). A carport's capacity and design logic sit closer to a power plant: engineered structures, often built and financed at commercial scale. But its injection mode, its ownership, and where it physically sits — a supermarket parking lot, a factory yard — sit closer to rooftop or commercial PV. It doesn't sit on either side of the roof/ground line; it sits on a continuum the line doesn't account for. And it's a useful reminder that capacity or surface area, the two variables everyone reaches for first, aren't always the right axis to classify on.
</p>

<p>
  Then there's the question of what "a plant" even is. The intuitive answer is: a set of arrays, physically grouped, doing the same thing. Fine — but now ask what "a connection" is. Several arrays can sit behind a single grid connection point. Or the same physical site can end up behind two or three separate connection points, depending on how and when it was built out — an extension added a year later, metered separately because that's how the paperwork happened to work at the time. So you can have one connection representing one plant, or three connections representing what is, on the ground, a single plant. Same physical reality, different administrative count, depending entirely on build history that has nothing to do with the electricity being produced.
</p>

<p>
  Layer onto that everything else that compounds before any comparison even starts:
</p>

<ul>
  <li>injection modality — total feed-in versus self-consumption with partial feed-in, which can turn one physical system into two distinct administrative records;</li>
  <li>installed capacity versus connected (rated) capacity, which routinely differ;</li>
  <li><strong>plug-and-play systems</strong> — plugged into a wall outlet, producing real electricity, but not connected in the formal sense a grid operator's registry expects. Plugged to the building. Invisible to the connection registry.</li>
</ul>

<p>
  And we haven't even gotten to what happens <em>after</em> a system is counted — production estimates, data sharing constraints, the fact that some of this is sensitive or simply private by design, no point de livraison published, no address disclosed. We've merely identified the system. We haven't started doing anything with it. It's worth naming honestly: this isn't a tidy taxonomy waiting for the right specialist to sort it out. It's a cheerful mess.
</p>

<h3>2. Three representations</h3>

<p>
  Each channel that tries to describe PV builds its own answer to the questions above — implicitly, usually without saying so. Before comparing them, it's worth laying out what each one is actually built to see, and what it structurally cannot.
</p>

<p>
  <strong>Grid connection data</strong> records an administrative event: a system officially connected to the network, tied to a point de livraison. It captures what has gone through the connection pipeline. It's strong on some attributes — capacity, connection date — precisely because those are what the paperwork requires. It's only as complete as that pipeline: truncation at low-reporting thresholds, registration lag, aggregation choices that hide small or recent systems.
</p>

<p>
  <strong>Open mapping</strong> works bottom-up: whoever surveys or digitizes a system tags it, at whatever granularity the local mapper chose, under a schema the community negotiates in the open — which is exactly what the proposal opening this piece is doing. Its strength is that it's not gated by any single operator or administrative process; a plug-and-play balcony unit can be tagged here even though it will never appear in a connection registry. Its weakness is the mirror image of that strength: coverage is only as good as who showed up to map it, and it's uneven by construction.
</p>

<p>
  <strong>Remote sensing</strong> (DeepPVMapper-type pipelines) sees a physical, visual object: pixels classified as PV, grouped into a cluster. It's bound to the date of the imagery — a system installed after the flight doesn't exist as far as the model is concerned — and to whatever the detection and clustering logic decides counts as "one" installation. Worth flagging here, because it matters later: remote sensing is, by construction, exhaustive over its coverage area and globally interoperable — the same method applied to imagery anywhere gives a comparable output, with no dependency on a public operator's registry existing or being any good.
</p>

<p>
  Three honest, internally consistent descriptions. Three different objects.
</p>

<h3>3. The comparison problem</h3>

<p>
  Given all that, why compare these sources at all? Because each one's blind spot is close to being another one's specialty. Grid data is authoritative on what has been formally connected and when, but blind to everything that hasn't gone through that process. Remote sensing is blind to intent, ownership, and anything installed after the last flight, but doesn't care whether an administrative process ever happened. Open mapping can catch what neither of the other two are built to catch, at the cost of even coverage.
</p>

<p>
  The scale of the mismatch isn't a rounding error. Two examples make that concrete.
</p>

<p>
  <strong>Pakistan</strong> had imported over 55.7 GW of solar panels cumulatively by the end of May 2026 [1]. Officially net-metered installed capacity, as reported by the distribution companies that are supposed to be tracking it, stood at just 6.5 GW at the close of the previous fiscal year — meaning at least 86% of everything imported never showed up as a formal, grid-tied connection [1]. Some of that gap is inventory sitting in warehouses or installation lag, but a meaningful share is systems that were never meant to touch the registry at all: commercial and industrial self-consumption setups, hybrid inverters paired with batteries, off-grid installations in areas net metering doesn't reach [1]. By 2025, distributed solar — net-metered, behind-the-meter, and off-grid combined — was already estimated to generate the equivalent of nearly half of all grid-supplied electricity in the country, almost entirely outside what the official connection data records [2].
</p>

<p>
  <strong>Germany</strong> shows the same pattern at a different scale. As of early 2026, just over 1.3 million plug-and-play balcony systems ("Balkonkraftwerke") were formally registered in the Bundesnetzagentur's Marktstammdatenregister [3]. But researchers at HTW Berlin, working from retailer and user surveys, estimated the real number of installed units at somewhere between 1.5 and 4 million as early as February 2025 — when only around 862,000 were officially on the books [4]. Put differently: independent estimates suggest something like a third to half of these systems ever get registered at all, even though registration has been simplified to a fifteen-minute online form since 2024 [4].
</p>

<p>
  In both cases, the grid operator isn't wrong about what it reports. It's reporting what its process is built to capture. The gap is structural, not a data quality problem waiting to be cleaned up — and it's large enough, in both directions, that treating any single source as ground truth would be a mistake with real consequences for grid planning and policy, not just an academic footnote.
</p>

<h3>4. Toward a shared ontology</h3>

<p>
  None of this argues for a single, unified data source that replaces the others — that's not realistic,
  and it isn't the goal. Grid operators need what grid operators need; remote sensing will always be
  bound to imagery dates; open mapping will always depend on who shows up.
</p>

<p>
  But if the goal is to avoid double-counting, to actually compare what these sources see,
  and to know when a gap reflects real underreporting rather than definitional mismatch, then
  some shared reference is worth having — not a perfect one, but one that names the ambiguities
  from §1 explicitly instead of letting each pipeline resolve them silently in its own conventions.
</p>

<p>
  An openly negotiated, imperfect ontology doesn't resolve the mess. What it does is give each modality —
  grid connection, remote sensing, crowdsourced mapping — a common surface to plug into, so that when
  they disagree, the disagreement is legible instead of just noise. This is where the MapYourGrid
  proposal earns a second mention. The current revision of the power plants descriptions is
  accessible <a href="https://wiki.openstreetmap.org/wiki/Proposal:Power_generation_storage">here</a>.
</p>

<p>
  The point was never to find the one true description of the elephant. It's to let everyone standing around it compare notes.
</p>

<h3>Sources</h3>

<ul>
  <li>[1] Business Recorder — Pakistan's solar revolution: phase-two in full swing, July 2026.</li>
  <li>[2] Ember — The solarisation of Pakistan's energy economy, June 2026.</li>
  <li>[3] 42watt.de — Balkonkraftwerk anmelden 2026: Anleitung für das Marktstammdatenregister, July 2026.</li>
  <li>[4] SOLARPUNK — Krasse Dunkelziffer: So viele Balkonkraftwerke gibt es in Deutschland, November 2025.</li>
</ul>]]></content><author><name>Gabriel Kasmi</name></author><category term="Open Data" /><category term="Photovoltaics" /><category term="OpenStreetMap" /><summary type="html"><![CDATA[Reconciling grid connection data, remote sensing, and OpenStreetMap: why comparing PV maps means agreeing on what a PV installation even is.]]></summary></entry><entry><title type="html">What Does a Solar Panel Detector Actually See?</title><link href="https://gabrielkasmi.github.io/blog/2026/07/13/wavelets-and-the-grid-detector/" rel="alternate" type="text/html" title="What Does a Solar Panel Detector Actually See?" /><published>2026-07-13T00:00:00+00:00</published><updated>2026-07-13T00:00:00+00:00</updated><id>https://gabrielkasmi.github.io/blog/2026/07/13/wavelets-and-the-grid-detector</id><content type="html" xml:base="https://gabrielkasmi.github.io/blog/2026/07/13/wavelets-and-the-grid-detector/"><![CDATA[<p>
  When I was developing <a href="/projects/#deeppvmapper">DeepPVMapper</a> — a deep learning
  pipeline that maps rooftop photovoltaic installations from aerial imagery across France — the
  question I got asked most wasn't how accurate the model was. It was simpler, and harder:
  does it actually detect solar panels?
</p>

<p>
  At first this sounds like the same question. It isn't. A model can score well on a held-out
  test set and still be latching onto the wrong thing. Precision and recall tell you the model
  agrees with your labels; they don't tell you why. And "why" is exactly what you need before
  trusting a model deployed at the scale of an entire country, because the failure modes that
  matter — the ones that surface once you leave the validation set and hit a new region, a new
  roof material, a new camera — are invisible to an aggregate metric.
</p>

<h3>Why "what does it see" is hard</h3>

<p>
  The instinctive answer is to reach for an explainability method, get a saliency map, and look
  at which pixels the model attended to. I did that. It's the standard toolkit — Grad-CAM,
  integrated gradients, occlusion-based attributions — and all of these answer the same question:
  <em>where</em> did the model look?
</p>

<p>
  That's useful, but it isn't the question I needed answered. For a model, a solar panel is not
  a semantic category — it's a distribution of pixel intensities. What I needed to know wasn't
  where a prediction came from on the image, but <em>what kind of visual structure</em> it was
  reacting to at that location. Was it the fine, repetitive texture of individual PV cells and
  their grid lines? The coarse rectangular outline of the array against the roof? Something else
  entirely — a skylight, a water tank, a shadow with the right aspect ratio? None of the existing
  pixel-domain tools could answer that. They could point at a location; they couldn't describe
  what was there from the model's point of view.
</p>

<h3>Decomposing the image into scales</h3>

<p>
  This gap is what pushed us to build a new attribution method based on <strong>wavelet
  decomposition</strong>, first introduced as the Wavelet Scale Attribution Method (WCAM), and
  later generalized as the
  <a href="https://gabrielkasmi.github.io/wam/">Wavelet Attribution Method (WAM)</a> at ICML 2025.
</p>

<p>
  Wavelets decompose an image into a set of scales — from coarse, low-frequency components (the
  general shape and layout of an object) to fine, high-frequency components (texture, edges,
  small repeated patterns). Unlike a plain pixel-domain heatmap, a wavelet attribution tells you,
  for a given location, <em>which scale</em> of visual information the model's decision actually
  depends on.
</p>

<figure class="blog-figure">
  <img src="/blog/img/wavelets-and-the-grid-detector/space-scale-fig6-gradcam-vs-wcam.png" alt="Classic Grad-CAM heatmap next to the WCAM decomposition of the same prediction into scale bands" />
  <figcaption>
    <strong>Figure 1.</strong> A classic pixel-domain saliency map (left) tells you <em>where</em>
    the model looked, but collapses every scale into a single heatmap. The WCAM (right) breaks the
    same prediction down by scale band (from coarse structures, &gt;8 px, to fine texture, 1–2 px),
    showing exactly which spatial frequency the decision relies on. Source:
    <a href="https://doi.org/10.1017/eds.2025.13">Kasmi et al., Environmental Data Science (2025)</a>.
  </figcaption>
</figure>

<p>
  Applied to a PV detection, this distinction is exactly what I was missing: does the model rely
  mostly on the coarse shape of the array — a dark rectangle on a lighter roof — or on the fine
  texture of individual modules and their grid pattern? Two models can produce an identical
  bounding box and an identical saliency heatmap while relying on completely different, and not
  equally reliable, visual evidence. This matters especially for PV systems, which are themselves
  multi-scale objects: the same installation can be described as a roof-sized system, a cluster of
  modules, or a few centimeters of grid line, depending on where you zoom in.
</p>

<figure class="blog-figure">
  <img src="/blog/img/wavelets-and-the-grid-detector/space-scale-fig3-pv-multiscale.png" alt="A rooftop PV system decomposed from the overall system down to the fine details of a single module" />
  <figcaption>
    <strong>Figure 2.</strong> The same PV installation, read at different scales: the overall
    system on the roof (~10 m), the array as a whole (~2.5 m), a group of modules (~1–2 m), a
    single module (&lt;1 m), and the fine grid pattern within a module (~0.1–0.2 m). A model can
    latch onto any one of these — and only some of them are reliably specific to solar panels.
    Source: <a href="https://doi.org/10.1017/eds.2025.13">Kasmi et al., Environmental Data Science (2025)</a>.
  </figcaption>
</figure>

<p>
  The clearest way to see what a model is actually keying on is to progressively strip away the
  wavelet components it considers least important and watch what survives.
</p>

<figure class="blog-figure">
  <img src="/blog/img/wavelets-and-the-grid-detector/grid-disappearance.gif" alt="Animation showing important zones and important components collapsing onto the grid pattern of a PV panel as less relevant wavelet components are removed" />
  <figcaption>
    <strong>Figure 3.</strong> Input image, the model's important zones (a standard pixel-domain
    heatmap), and the important components once the prediction is read in the wavelet domain. As
    less relevant components are removed, what's left is the grid — not the panel as a whole.
    Source: <a href="https://theconversation.com/photovolta-que-et-reseau-electrique-comment-une-ia-fiable-et-transparente-pourrait-faciliter-la-decarbonation-261681">Kasmi, The Conversation (2025)</a>.
  </figcaption>
</figure>

<h3>The reveal: an excellent grid detector</h3>

<p>
  Running this analysis at scale across DeepPVMapper's predictions gave a clear, and slightly
  humbling, answer. The model wasn't, in general, detecting solar panels. It was detecting
  <strong>grids</strong> — regular, high-frequency repeated patterns at a particular scale.
  Actual PV arrays have that pattern, which is why the model worked as well as it did. But so do
  a number of other things: greenhouse roofs, certain skylights, corrugated roofing, solar water
  heaters, parking lot shade structures. Wherever that pattern showed up, at the scale the model
  had learned to key on, it fired — regardless of whether a solar panel was actually there.
</p>

<div class="blog-figure-row">
  <figure class="blog-figure">
    <img src="/blog/img/wavelets-and-the-grid-detector/false-positive-farm-building.png" alt="False positive detection on a corrugated farm building roof, flagged as a PV array" />
    <figcaption>
      <strong>Figure 4a.</strong> A farm building roof in the Manche, flagged as a PV array. The
      corrugated roofing produces the same regular, high-frequency pattern the model associates
      with solar panels.
    </figcaption>
  </figure>
  <figure class="blog-figure">
    <img src="/blog/img/wavelets-and-the-grid-detector/false-positive-greenhouse.png" alt="False positive detection on greenhouse tunnels, flagged as a PV array" />
    <figcaption>
      <strong>Figure 4b.</strong> Greenhouse tunnels, also flagged. Same story: a grid-like texture
      at the scale the model has learned to key on. Buildings like these were common enough in the
      Manche to produce a steady stream of false positives.
    </figcaption>
  </figure>
</div>

<p>
  Tellingly, a genuine PV installation is picked up for the same reason — not because the model
  recognizes "solar panel" as a category, but because it finds the same grid signature.
</p>

<figure class="blog-figure">
  <img src="/blog/img/wavelets-and-the-grid-detector/true-positive-pv-panel.png" alt="True positive detection on a small rooftop PV installation" />
  <figcaption>
    <strong>Figure 5.</strong> A correctly detected rooftop PV installation. Its wavelet-scale
    signature — a regular grid at a specific scale — is what the model actually relies on here too,
    which is precisely why it can't reliably tell this apart from Figures 4a and 4b.
  </figcaption>
</figure>

<p>
  This is a far more actionable diagnosis than "false positive rate is X%." It names the specific
  visual confound driving the errors, which means it can be acted on directly: targeted
  hard-negative mining on grid-textured non-PV structures, augmentation that decouples grid
  texture from array shape, or an architecture change that forces the model to weigh coarse-scale
  shape cues rather than fine-scale texture alone.
</p>
<p>
  Circling back to accuracy, this pattern also helps interpret the sharp precision differences that
  we have observed across France. As many corrugated roofing structures appear in rural areas in the North West
  of France and especially in the Manche département, these roofs triggered a lot of false positives,
  driving the precision down. Similarly, the fact that well defined gridded PV systems are less present
  in Brittany also helps explain why the false negative rate was higher in locations
  such as Morbihan or Côtes-d'Armor.
</p>

<h3>Why this matters beyond one model</h3>

<p>
  None of this is specific to solar panels. Any CNN-based detector trained on a finite set of
  positive examples can end up encoding a proxy feature — a texture, a repeated pattern, a color —
  rather than the semantic category it's supposed to represent. Wavelet-scale attribution gives a
  way to check, for a specific model and a specific prediction, which one it actually learned —
  before that gap costs accuracy on data you haven't seen yet.
</p>

<p>
  This example is also a reminder that the failure mode is a property of the training data, not
  just the architecture. The patterns a model relies on at deployment are the patterns it was
  shown during training — change what's in the training set, and different patterns can take over.
  A natural fix, already explored by
  <a href="https://www.sciencedirect.com/science/article/pii/S0306261925003605">Thébaud <i>et al.</i></a>,
  is to explicitly train on a wider range of rooftop types, and it does improve detection on
  previously mismatched building types.
</p>

<p>
  But that result raises a question I don't think can be answered without checking: is the gain
  coming from genuine diversity — the model learning a richer, more invariant notion of "solar
  panel" — or is it a more mundane form of hard-negative mining in disguise? Adding new regions to
  a training set doesn't teach a CNN anything about geography; it only changes which patterns show
  up as positives and negatives. If those new regions happen to contain grid-textured non-PV
  structures the original set didn't, then what looks like an improvement from geographic diversity
  is really the same shortcut being patched with more examples of what to exclude — not a
  different, more reliable feature.
</p>

<p>
  The distinction matters because it predicts different behavior down the line. A model that has
  genuinely stopped keying on the grid pattern should generalize to a new, unseen texture confound.
  A model that has simply seen more instances of "grid, but not PV" will most likely fail again the
  next time a genuinely novel one shows up. Wavelet-scale attribution gives a direct way to tell
  the two apart — check whether the retrained model's true and false positives still share the same
  fine-scale signature, or whether it has started relying on coarser, shape-based cues instead.
  That's the natural next step here, and one I'd rather test than assume.
</p>

<h3>Further reading</h3>

<ul>
  <li>
    [In French] A longer, less technical version of this argument, on how reliable and transparent
    AI can support the integration of solar power into the grid, published in
    <a href="https://theconversation.com/photovolta-que-et-reseau-electrique-comment-une-ia-fiable-et-transparente-pourrait-faciliter-la-decarbonation-261681">The Conversation</a>.
  </li>
  <li>
    The full method behind this analysis, applied specifically to rooftop PV mapping: <em>Space-scale
    exploration of the poor reliability of deep learning models: the case of the remote sensing of
    rooftop photovoltaic systems</em>, published in
    <a href="https://doi.org/10.1017/eds.2025.13">Environmental Data Science</a>.
  </li>
  <li>
    The general-purpose version of the method: <em>One Wave To Explain Them All: A Unifying
    Perspective on Feature Attribution</em>, published at
    <a href="https://proceedings.mlr.press/v267/kasmi25a.html">ICML 2025</a>
    (<a href="https://gabrielkasmi.github.io/wam/">project page</a>).
  </li>
</ul>]]></content><author><name>Gabriel Kasmi</name></author><category term="Explainable AI" /><category term="Computer Vision" /><category term="DeepPVMapper" /><summary type="html"><![CDATA[Using wavelet decomposition to explain what a deep learning model actually detects, and why DeepPVMapper turned out to be a grid detector.]]></summary></entry></feed>