Introduction
Field engineers operating in remote substations frequently encounter connectivity limitations when evaluating the safety of proposed grid adjustments. While physics-informed neural surrogates for AC Optimal Power Flow can deliver these assessments in milliseconds, they typically rely on cloud infrastructure. Beyond the risk of complete service loss in low-connectivity environments, transmitting telemetry off-site introduces critical data privacy and security vulnerabilities for infrastructure operators.
This project addresses these operational constraints by exploring the limits of model compression. Specifically, it evaluates how significantly a state-of-the-art power grid surrogate can be quantized without compromising safety thresholds, while integrating the optimized model into a localized small language model (SLM) assistant capable of executing entirely offline on standard client hardware.
About me
My name is Margaux Vallet, and I am currently finishing my Master’s degree at Imperial College London in Applied Computational Science and Engineering. For my Masters’ research project, I collaborated with Microsoft to create an agent with Foundry Local for field engineers in electricity grids. The agent consists of an SLM able to call GridSFM, a neural surrogate for grid optimization, and a RAG layer, to assist field engineers in their daily tasks.
I would like to thank Lee Stott who was my supervisor during this project.
Research Contributions
- A systematic quantization study of GridSFM across 53 power grid topologies, including synthetic (PGLib-OPF) [4] and real-world (MSR corpus) network structures, at three precision levels (FP16 naive, FP16 mixed precision, INT8) evaluated against a FP32 baseline.
- Identification of a topology-specific failure mode in INT8 quantization, showing that degradation is not uniformly distributed but concentrated in a specific family of grid topologies.
- A methodology for evaluating whether a language-model interface stays grounded in the output of the numerical tool it wraps, applicable to any agentic system built on top of a deterministic predictive model.
- A safety-oriented agent architecture that structurally separates numerical inference from language reasoning, which enables to attribute errors independently to the predictive model or to the language model’s interpretation of it, rather than treating “the agent was wrong” as an undifferentiated failure category.
Project Overview
AC Optimal Power Flow (AC-OPF) sits at the centre of day-to-day grid operations. It aims to minimize the generation cost while respecting the physical constraints and the AC non-linear equations. It is also very expensive to compute: traditional solvers using interior-point methods can take minutes to hours per run, and grid operators need answers many times a day, under strict reliability deadlines. The stakes are high: according to estimates, the limitations of standard methods degrade AC-OPF quality, which could result in $20B per year in congestion losses and multi-terrawatt-hour renewable curtailments.
To address theses issues, Microsoft Research recently released GridSFM [1], a physics-informed neural surrogate that computes a full AC-OPF operating point (voltages, generator dispatch, branch flows, and a feasibility assessment) in milliseconds, and, generalizes across grid topologies rather than being locked to one network, contrary to most existing models. This tackles a limitation common to earlier physics-informed and graph-based OPF surrogates [2], which typically generalize poorly beyond the topology they were trained on.
However, GridSFM has never been deployed locally on edge devices or given a natural language interface. That is the gap this project aims to close, by creating a field engineer’s assistant with three parts:
- GridSFM as the numerical core, that we quantize to see how small it can go before it stops being trustworthy.
- A small language model (SLM) as the natural-language interface, running on-device via Microsoft Foundry Local. Recent work suggests carefully deployed SLMs can match much larger models on domain-specific tasks [3], which is the premise this project puts under empirical test.
- A Retrieval-Augmented Generation (RAG) layer, grounding the assistant’s answers in real technical documentation.
The objective was not just to make it smaller, but to find the minimum viable precision at which GridSFM remains safe for decision support, and to check that the language model sitting in front of it does not quietly undermine that safety by making things up.
Project Journey
To address these constraints, the project was organized into two complementary core objectives: applying quantization techniques to compress GridSFM, and designing a safety-oriented agent architecture capable of reliable offline deployment.
Setting the baseline
The experimental baseline was established using the default FP32 model checkpoint (gridsfm_open_v1.1.pt) evaluated across 53 power grid topologies from the PGLib-OPF benchmark and Microsoft’s MSR corpus, covering networks ranging from 500 to nearly 3,900 buses. Ground-truth reference values were generated by solving each topology via the IPOPT non-linear solver (with a convergence tolerance of 1e- and an acceptable tolerance of ). This enabled direct comparative analysis of every quantized variant against both FP32 model parity and true AC-OPF optimal feasibility.
Diagnosing failures in FP16
Direct naive conversion to FP16 half-precision induced severe degradation in feasibility prediction accuracy across all evaluation topologies. Diagnostic isolation initially targeted the scatter_mean and scatter_reduce operations within GridSFM’s FusionLayer. Because these pooling operations lack explicit autocast policies, they silently inherit FP16 precision from upstream tensor operations. However, applying explicit autocast management failed to resolve the issue. Enforcing standard FP32 precision specifically for the pooling operations while keeping the remainder of the network in FP16 similarly failed to restore stability.
Inspection of the intermediate tensors entering FusionLayer revealed that NaN values were propagating from earlier operations prior to entering the pooling layer, where a downstream _nan_to_zero call masked their presence. Rather than yielding an immediate mitigation, these experiments functioned as a systematic elimination process, showing that establishing a plausible domain for a numerical defect requires empirical verification before confirming its root cause.
Making INT8 work on Apple Silicon
Dynamic INT8 quantization of the nn.Linear layers (roughly 89.3% of GridSFM’s parameters) yielded greater stability after addressing a platform-specific dependency: PyTorch’s native quantize_dynamic required explicitly setting the quantization backend to qnnpack for Apple Silicon execution.
Building the agent
Foundry Local’s SDK is structured for generative models and lacks native support for running non-generative regression architectures like GridSFM within its runtime. Decoupling GridSFM into a standalone FastAPI service callable by the SLM provided a key architectural advantage: standardizing the interface allowed for a clear operational separation between numerical prediction accuracy and language model interpretation accuracy, isolating two distinct failure modes.
The following is an example of prediction given by the GridSFM tool for the msr_texas grid:
{
“success”: true,
“duration_ms”: 2627.44,
“outputs”: {
“bus_predictions”: [
[0.1279, 1.0500],
[0.1538, 1.0486],
[0.2088, 1.0418],
// … 3,886 more buses
],
“generator_predictions”: [
[0.2380, -0.0190],
[0.0000, -0.1789],
[0.0000, -0.0868],
// … 506 more generators
],
“feasibility”: 0.999988
},
“message”: null
}
Experimental Protocol
For readers looking to replicate or build on this work, here are the concrete details of how the evaluation was run.
- Runs per topology. Latency was measured over 50 timed iterations per topology per precision variant, after 5 discarded warm-up iterations, with mean and standard deviation reported. Ground truth came from a single deterministic IPOPT solve per topology, since the solver itself is not stochastic, so no seed averaging was needed on that side.
- Metrics used. Beyond feasibility delta (the absolute difference in predicted feasibility probability between FP32 and a quantized variant), the evaluation tracked bus-level and generator-level MAE and RMSE, computed both against the FP32 baseline (parity) and against the IPOPT ground truth (accuracy), plus maximum absolute error per topology and on-disk model size.
- How agent faithfulness was annotated. 50 tool-grounded agent responses were manually classified into one of four categories: grounded (accurately represents the tool output with no unsupported additions), reasonable gloss (paraphrases or rounds the tool output without changing its meaning), fabricated (introduces a numerical value or judgement the tool’s JSON output does not support or contradicts), or non-substantive (a refusal or a request for clarification). Annotation was performed by a single rater; as a limited consistency check, two pairs of identical queries were included in the batch and annotated blind to the fact that they were repeats, and both pairs received consistent labels. A second, independent annotator with a formal agreement measure (e.g. Cohen’s kappa) is a natural next step and is listed under Future Development.
Technical Details
- Model: GridSFM-Open (Microsoft Research), loaded from the official Hugging Face checkpoint, with inference code from the public GridSFM repository.
- Environment: Python 3.11, conda, PyTorch 2.8.0, pinned explicitly because quantization behavior can shift across PyTorch releases.
- Hardware: Apple Silicon, CPU-only inference throughout, reflecting the intended field-deployment scenario where a dedicated GPU can’t be assumed. Quantization backend: QNNPACK, for ARM compatibility.
- Quantization: INT8 dynamic quantization of nn.Linear layers via torch.ao.quantization; FP16 via .half() and CPU torch.autocast.
- Serving layer: GridSFM wrapped in a FastAPI service with a fixed JSON schema (bus_predictions, generator_predictions, feasibility), called as a tool by the language model rather than loaded inside the model runtime.
- Local LLM runtime: Microsoft Foundry Local, running qwen2.5-7b as the primary tool-using SLM, with qwen2.5-0.5b used as a smaller comparison point.
- Grounding: A RAG layer over GridSFM documentation and grid technical references, to stop the agent inventing descriptions of things it hasn’t actually looked up.
- Verification: Every quantized checkpoint was loaded with strict=True so mismatched parameters fail loudly rather than silently, and each variant was tested before and after serialization to confirm saving and reloading didn’t change its behavior.
Results and Outcomes
Quantization
Across all 53 grid topologies, INT8 dynamic quantization maintained high accuracy across most of the dataset, with 37 topologies (70%) showing a feasibility prediction delta of less than 0.01 compared to the FP32 baseline. The failures that did occur were concentrated rather than uniformly distributed. The seven PGLib-OPF cases derived from the Polish transmission grid exhibited a mean feasibility delta of 0.667, compared to 0.037 across all remaining topologies. These specific networks are well-documented in Optimal Power Flow (OPF) literature as being numerically ill-conditioned even for classical deterministic solvers [5], which is a reassuring, if partial, explanation.
Due to PyTorch runtime overheads and selective quantization of nn.Linear layers, INT8 demonstrated the following trade-offs relative to FP32: INT8 also ran 1.07 to 1.36 times slower than FP32, and shrank the checkpoint by 2.87 times rather than the theoretical 4 times.
FP16, meanwhile, was simply broken: 38 of 53 topologies had their feasibility classification completely inverted.
In conclusion, for this architecture, model compression did not mean performance acceleration. The highest degree of quantization also produced unpredictable outputs, with edge-case degradation concentrated within a specific topology family rather than spreading evenly across the evaluation suite.
Agent faithfulness
Evaluating model safety required looking beyond numerical precision to analyze how faithfully the language model (SLM) interpreted output from the underlying numerical solver. Manual annotation of 50 agent responses generated by the 7B parameter model against raw tool outputs gave the following breakdown: 60% grounded, 16% reasonable gloss, 20% non-substantive, and 2% fabricated.
While a 2% fabrication rate appears promising, evaluating the identical query set on a smaller model (qwen2.5-0.5b) demonstrated a steep decline in fidelity: grounded responses dropped from 60% to 10%, while fabrications increased sharply from 2% to 48%. Thus, model scale directly impacts systemic safety in edge deployments. Larger parameter counts do not merely improve narrative fluency; they serve as a risk-mitigation mechanism to ensure the agent remains faithful to the underlying deterministic model.
Lessons learned
- Quantization failures can be topology-dependent: An aggregate mean INT8 feasibility delta of 0.120 suggested moderate and acceptable degradation. However, disaggregated analysis revealed near-perfect performance across 70% of network topologies alongside severe failure concentrated within a single topology family. Evaluating models purely via global performance metrics masks localized failure modes.
- Downstream sanitization can mask precision instabilities: The presence of a _nan_to_zero operation enabled the model to output plausible predictions long after NaN values had corrupted internal state tensors. This highlights the risk of silent numerical degradation when error-handling operations obscure upstream precision failures.
- Language model scale directly drives systemic safety: Our experiments revealed that SLM scale directly impacts response fidelity. Reducing parameter size caused fabrication rates to increase from 2% to 48%, proving that a numerically precise backend cannot guarantee a reliable language interface.
- Architectural decoupling can be safety-critical: Isolating numerical inference behind a fixed-schema API enabled independent verification of the predictive engine and the language interface. This architecture establishes that numerical correctness and semantic fidelity represent distinct failure domains requiring isolated evaluation frameworks.
Implications for Teaching & Research
This project highlights several key architectural and methodological patterns that extend beyond GridSFM, offering valuable case studies for computer science, machine learning, and domain-specific AI curricula.
- Quantization Dynamics in Graph-Based vs. Transformer Architectures: Unlike standard Transformer-based LLMs (which often tolerate low-bit quantization smoothly due to localized attention mechanisms), graph-based neural surrogates like GridSFM exhibit distinct sensitivities. As message-passing operations continuously aggregate features across connected grid topologies, precision losses accumulate non-linearly over graph edges. This contrast highlights how structural inductive biases alter numerical compression limits compared to sequence-based models.
- Instructional Case Study on Numerical Instability: The FP16 investigation serves as a practical lesson in debugging complex ML pipelines. It illustrates how numerical failures can persist through standard remediation attempts (such as autocast integration and targeted FP32 pooling) until intermediate tensor states are explicitly inspected, countering the assumption that quantization is a purely mechanical procedure.
- Topology-Dependent Robustness and Benchmark Design: The concentrated failure rate observed within the Polish grid topologies demonstrates how global performance averages hide structural vulnerabilities. This case provides a clear framework for teaching benchmark design, stressing the need for disaggregated evaluations when deploying models to safety-critical environments.
- Architectural Decoupling for AI Safety: By isolating the deterministic numerical tool from the natural language interface via a fixed-schema API, the architecture establishes a model for trustworthy AI system design. This approach allows students and researchers to independently evaluate numerical correctness and semantic fidelity, preventing separate failure modes from collapsing into a single error state.
- Reproducible Edge AI Framework for Student Engineering: The integration of localized runtime environments, quantized domain models, and structured API boundaries offers a practical blueprint for offline decision-support systems across privacy-sensitive domains.
Future Development
A few threads are left deliberately open for what comes next:
- Tracing the FP16 failure to its source. The investigation narrowed the NaN’s origin to the computation upstream of FusionLayer, but didn’t pin down the exact operation. Tracing it through the attention and message-passing stack would be a natural next step.
- A connectivity analysis of the Polish grid family. The INT8 failures cluster there, but without computing structural metrics like node degree, meshedness, or branch-to-bus ratio across the dataset, it remains a hypothesis rather than a confirmed explanation. Answering this would clarify whether the failure mode generalizes to other real-world grids or is specific to this network.
- A formal operational acceptance threshold, developed with input from grid operators, against which quantization-induced dispatch errors could actually be judged acceptable or not, an operational threshold this project deliberately refrained from defining unilaterally.
- Broader agent evaluation, extending faithfulness testing across more SLM sizes and families, a second independent annotator with a formal agreement measure, and GPU-based mixed precision as a comparison to the CPU-only results here.
Conclusion
This project set out to determine whether a state-of-the-art power grid surrogate could be compressed for fully offline deployment without compromising the strict safety requirements of critical infrastructure. The findings reveal a nuanced reality:
- INT8 quantization is generally viable across most scenarios, yet exhibits severe degradation within a specific, structurally identifiable family of grid topologies.
- FP16 is unsafe in its current implementation, requiring further architectural adjustments to prevent silent tensor corruption.
- A numerically accurate backend model does not guarantee a safe user interface; the language model layer introduces independent risk factors that decrease with parameter scale but are never fully eliminated.
Achieving trustworthy local AI for critical infrastructure requires rigorous evaluation across both operational layers, the underlying numerical engine and the natural language interface, rather than relying solely on isolated benchmarking metrics.
Final note
If you are interested in building offline, on-device AI systems of your own, these Microsoft resources are a good place to start:
- Microsoft Foundry Local for running language models fully on-device without cloud dependency.
- A tutorial on how to add a service app as a tool in a Foundry agent
- The discord channel for Microsoft Foundry, which is very helpful
Lastly, I would be happy to connect if you would like to know more, or have thoughts on this project. Feel free to reach out to me on LinkedIn.
References
[1] Yang, W., Britto Mattos Lima, A., Spina, T. V., Fowers, S., Zhang, B., White, C. GridSFM: A Foundation Model for AC Optimal Power Flow. Microsoft Research, 2026.
[2] Al-Ismail, F. S. Physics Informed Neural Networks-Based AC Optimal Power Flow Under High RES Penetration.IEEE Access, 12:189297 to 189306, 2024.
[3] Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., Lane, N. D., Xu, M. Small Language Models: Survey, Measurements, and Insights. 2025.
[4] Babaeinejadsarookolaee, S. et al. The Power Grid Library for Benchmarking AC Optimal Power Flow Algorithms.2021.
[5] Snodgrass, J., Kunkolienkar, S., Habiba, U., Liu, Y., Stevens, M., Safdarian, F., Overbye, T., Korab, R. Case Study of Enhancing the MATPOWER Polish Electric Grid. 2022 IEEE Texas Power and Energy Conference (TPEC), College Station, TX, February 2022.

