blubanyan logo

Best Data Sources For Solar Yield Models

Best Data Sources For Solar Yield Models

A weak solar yield model can miss revenue by hundreds of thousands of dollars per year. If I’m modeling U.S. solar output, I’d use five data sources together: satellite/reanalysis data for the starting weather picture, on-site sensors for site truth, SCADA for plant output, inverter/string data for fault checks, and asset history for loss context.

Here’s the short version:

  • Satellite and reanalysis data help with long-term energy estimates and early-stage screening.
  • On-site weather sensors help correct weather bias at the plant.
  • SCADA data shows what the plant produced and when.
  • Inverter and string-level data helps find unit-level problems.
  • Asset and maintenance records explain downtime, repeat failures, and long-term drift.

A few numbers show why this matters:

  • A 2% to 3% annual forecast error on a 100 MW portfolio under a $0.05/kWh PPA can shift yearly revenue by a large amount.
  • U.S. utility-scale and C&I systems in one study underperformed modeled financing expectations by 6.3%.
  • Remote irradiance data can carry 3% to 5% error before plant-level correction.
  • In one NREL dataset, only 41% of 19,460 inverters passed data quality checks.

If I had to boil the whole article down to one point, it would be this: no single data source is enough. You need a stack that matches the job, whether that job is bankability, monthly revenue forecasting, curtailment tracking, or fault diagnosis.

What to compare across each source:

  • Coverage
  • Latency and resolution
  • Error risk
  • Fit for forecasting
5-Tier Solar Yield Model Data Stack: Sources, Uses & Error Risks
5-Tier Solar Yield Model Data Stack: Sources, Uses & Error Risks

Quick Comparison

Data sourceMain useMain issueBest fit
Satellite / reanalysis weatherLong-range resource baselineSite bias, coarse grids in some casesDevelopment, diligence, portfolio studies
On-site sensorsPlant-level weather truthSoiling, drift, bad placementBias correction, PR work, contract support
SCADAPlant output and statusMapping, unit, and timestamp errorsProduction tracking, curtailment, back-testing
Inverter / string dataComponent fault diagnosisVendor field mismatch, missing dataFailure analysis, degradation checks
Asset / maintenance recordsDowntime and lifecycle contextIncomplete logs, disconnected systemsAvailability, repeat failure risk, long-range loss assumptions

So if you’re asking, “What are the best data sources for solar yield models?” my answer is simple: start with remote weather data, correct it with site sensors, track plant output with SCADA, use inverter data to explain equipment loss, and use solar business management software to track maintenance records and explain what changed over time. That stack gives you the clearest path to cleaner forecasts and fewer surprises.

1. Satellite and Reanalysis Weather Data

Satellite and reanalysis data give solar yield models their broadest starting point: wide coverage, long history, and no site hardware required. That makes them a strong fit for early-stage development, M&A diligence, and portfolio screening when on-site sensors still aren’t in place. The tradeoff is pretty simple: they score high on coverage, but lower on site-level precision and live-use value.

Coverage

NSRDB (National Solar Radiation Database) is a common historical baseline for U.S. solar portfolios. Its historical product covers 1998 to present at 4 km x 4 km and 30-minute resolution [4]. Newer products provide 2 km, 5-minute data for the continental U.S., Hawaii, Mexico, and the Caribbean Islands, plus 2 km, 10-minute data for North and South America [4]. It includes GHI, DNI, and DHI, and its historical series are serially complete, which helps cut down gaps during model training [2].

ERA5 is a global reanalysis product covering 1940 to present at hourly intervals on a 0.25° x 0.25° grid. You get long, global coverage, but you give up some site-level detail. That makes ERA5 a good option when comparing assets across several countries or setting one standard baseline for long-range analysis.

Solcast and SolarAnywhere are commercial services built for solar forecasting. Solcast offers spatial resolution down to 90 meters, with temporal outputs from 5 to 60 minutes and forecast updates every 5 to 15 minutes [10]. SolarAnywhere offers spatial resolution down to 1 km and temporal resolution as fine as 1 minute.

Latency and Resolution

Latency decides where each dataset fits.

ERA5 and NSRDB are built for historical reconstruction, not live operations. NSRDB has about a 1-year lag on its historical products [4]. ERA5 is also geared toward looking backward. So if you’re handling intraday dispatch or battery scheduling, these aren’t the tools for the job.

Solcast and SolarAnywhere are aimed at operations. Solcast refreshes forecasts every 5 to 15 minutes and covers a window from 7 days back to 14 days ahead [10]. SolarAnywhere’s short-horizon products use cloud-motion-vector methods and refresh as new satellite imagery comes in, with forecast ranges from 1 minute to 5 hours ahead. For day-ahead and intraday decisions, these commercial feeds are the practical pick.

Error Risk

No satellite or reanalysis dataset is free of bias. Common error sources include cloud misclassification, terrain shading, snow cover, aerosols, humidity, and local weather patterns that coarse grids just can’t fully see. ERA5’s global mean irradiance bias has been reduced by 50% to 75% compared with earlier reanalyses, but it still smooths out site-level swings in complex terrain. Satellite-derived products often beat reanalysis for many solar use cases, but they can still miss fast-moving cloud edges and changing atmospheric conditions.

Here’s the key distinction:

  • Random error can often be reduced through blending and calibration.
  • Systematic bias needs local observations to correct.

A benchmark against 57 BSRN stations found measurable bias in all 10 datasets, which is a strong reminder that ground-truth calibration still matters [12].

Fit for Business Forecasting

DatasetBest Use CaseLimitation
NSRDBU.S. resource assessment, historical benchmarking, bankabilityAbout a 1-year historical lag; newer high-resolution products are regionally limited
ERA5Global benchmarking, long climate baselines, multi-country portfoliosCoarser grid; less precise in terrain-sensitive or localized sites
SolcastOperational forecasting, dispatch optimization, intraday schedulingLess suited to deep-history climatology than archive datasets
SolarAnywhereShort-horizon forecasting, fast-refresh operational workflowsProduct specifics vary by version and region

Use NSRDB or ERA5 for historical resource assessment, then validate with on-site sensors before using satellite data for bankable forecasting. For site-level calibration, compare these baselines with on-site weather stations and irradiance sensors.

2. On-Site Weather Stations and Irradiance Sensors

Once you have broad weather baselines, on-site sensors bring the model back to what’s happening at the plant. Satellite data gives you coverage across a big area. On-site instruments give you plant-level ground truth.

Coverage

The main tools here are thermopile pyranometers and PV reference cells. Pyranometers measure GHI and POA in W/m². Reference cells track module response more closely, which makes them useful for POA monitoring and performance ratio work.

The catch is coverage. In many plants, a single GHI sensor and a single POA sensor may stand in for a large array. That means differences in tilt, azimuth, and shading can slip through the cracks. Sensor placement matters too. If a unit sits near inverter skids or other obstructions, the readings can pick up bias.

Latency and Resolution

This is where on-site sensors have a clear edge over satellite products. A well-set-up MET station can sample irradiance every 1 to 5 seconds and log averaged values every 1 to 5 minutes, with data reaching the SCADA system within seconds to a few minutes [1].

That’s much faster than the 15- to 60-minute intervals common with satellite feeds. And that speed matters. It can catch ramp events, inverter clipping, and short cloud-driven swings that lower-frequency data would blur.

Error Risk

Most errors from on-site sensors come from installation and upkeep.

Soiling on pyranometer domes or reference cell surfaces can lower measured irradiance. Best-practice guidance says cleaning should happen often enough to keep degradation below 1% to 2% [14].

Misalignment is another common problem. GHI pyranometers need to be mounted perfectly level. Even small tilt errors can create cosine response bias. Reference cells bring their own issue set, mainly spectral and temperature sensitivity, so temperature correction helps line them up more closely with module behavior.

Calibration drift is slower, but it can do real damage. High-grade pyranometers should be recalibrated about every two years. One study that tracked 23 pyranometers found a dramatic improvement in RMSE after calibration [15]. That’s a strong reminder that uncalibrated sensors can skew yield-model inputs in a big way.

Fit for Business Forecasting

On-site sensors work best as a localization and validation layer, not as a standalone forecasting source. For short-term forecasting, from minutes to a few hours, high-frequency site data can support nowcasting. For day-ahead and longer horizons, it helps bias weather inputs.

It also helps tie irradiance data back to plant performance and reporting. Blu Banyan‘s SolarSuccess can centralize MET station data alongside performance monitoring and financial reporting.

InstrumentStrengthsKey Error RisksBest Fit
Class A PyranometerStandardized irradiance measurement for GHI and POASoiling, leveling errors, calibration driftUtility-scale forecasting, compliance, PR monitoring
PV Reference CellMatches module spectral and angular responseSpectral mismatch, temperature sensitivityPOA monitoring, performance diagnostics
Full MET StationCaptures ambient temperature, wind, module temperature, and irradianceSiting errors, data-mapping errorsPlant-level operations and anomaly detection

SCADA then turns these point measurements into plant-wide operating data.

3. SCADA System Data

SCADA takes site sensor data and turns it into plant-wide telemetry. A MET station gives you irradiance and temperature at one fixed spot. SCADA goes much further. It pulls data from inverters, combiner boxes, power meters, tracker controllers, and weather stations into one plant-level view.

That’s why SCADA is the main operating record for yield modeling. It’s the layer that turns weather and sensor inputs into outage data, curtailment data, and benchmark data. Put simply, SCADA sits between weather inputs and equipment-level diagnostics.

Coverage

In a utility-scale U.S. plant, SCADA often tracks hundreds to thousands of tags. These usually include active power, reactive power, AC voltage, equipment status, fault codes, curtailment signals, and site readings.

That broad view helps you see more than total energy production. You can also see why output looks the way it does. Is a unit faulted? Curtailed? Running as expected? SCADA helps answer that.

Coverage quality depends a lot on commissioning. Smaller commercial sites sometimes have only part of the full setup. They may skip string-level data or secondary sensors. And some tags that matter for modeling – like module back-of-panel temperature or detailed tracker status – may not be wired in at all.

Latency and Resolution

SCADA screens usually refresh every 1 to 5 seconds. Historians usually store data at 1- to 5-minute intervals, while API or integration lag often falls between 15 and 60 minutes [1].

That level of detail is usually enough for day-ahead forecasting.

Error Risk

This is the part where SCADA demands care. Common issues in U.S. solar plants include point mapping errors, communications gaps, scaling mistakes, and timestamp misalignment.

Point mapping errors can be especially hard to spot. On multi-vendor sites using Modbus, DNP3, or IEC 61850, a bad register map can assign one inverter’s data to the wrong unit or report power in MW when the model expects kW. Wrong point mapping or bad data typing can also assign one inverter’s data to the wrong unit.

Another problem shows up when a field device drops offline. Some SCADA historians keep the last value instead of recording a gap. On paper, that can make an offline inverter look like it’s still running. If you train a model on that history, the model gets the wrong story.

The table below separates data issues you can fix from actual production losses.

Error TypeWhat Goes WrongFix
Point mapping errorsWrong inverter data attributed to wrong unit; kW vs. kVA confusionValidate point lists against single-line diagrams at commissioning
Scaling mistakesPower logged in MW interpreted as kW; irradiance units mixedConfirm units against device ratings at commissioning
Timestamp driftSCADA and satellite data misaligned; models train on wrong slicesSync all devices to NTP; preserve device-side timestamps

Fit for Business Forecasting

SCADA is the core operating layer for yield modeling. But on its own, it’s not enough. It still needs weather inputs and equipment records to sort outages, curtailment, and actual performance losses.

When teams fix those gaps, the results can be hard to ignore. A U.S. independent power producer standardized tag names and units across 20+ plants, then reprocessed three years of SCADA history. That work narrowed its fleet-wide P50 forecast error from roughly 6% to 8% down to 3% to 4% relative to actuals, mainly by classifying curtailment and outage periods the right way [1].

Platforms like Blu Banyan’s SolarSuccess can help bring SCADA data together with asset records and financial reporting. For operators with a spread-out portfolio, that makes it easier to keep data pipelines consistent across sites.

Inverter logs and string-level data can catch faults that plant-wide SCADA may miss.

4. Inverter Logs and String-Level Data

SCADA tells you what the plant did. Inverter logs and string-level data help you see why it happened at the component level.

Inverter logs usually include AC output power, DC voltage and current for each MPPT input, internal temperatures, fault codes, and status signals. String-level monitoring goes one step further by recording current for each string, and sometimes voltage too. That extra detail matters. It lets you spot one weak string instead of chasing a vague array-wide problem. In short, this data is best for diagnosing loss, not for estimating weather.

Coverage

Older or smaller inverters often give you only rolled-up DC and AC values, with no per-string detail. At utility-scale sites with string monitoring hardware, operators can track each string and catch local issues such as soiling, shading, MC4 connector corrosion, or bypass diode failures by comparing parallel strings under the same irradiance conditions.

Inverter logs offer almost nothing on site conditions. There’s no irradiance reading, no ambient temperature, and no wind speed. So if you want to model yield, you need to pair inverter data with weather inputs.

That said, coverage only helps if the data is clean enough to compare across strings and inverters.

Latency and Resolution

Inverter data usually comes through an API or SCADA system at a 1- to 15-minute cadence. Some high-end setups support sub-minute sampling for diagnostic campaigns, but for long-range forecasting, 5- to 15-minute data is usually enough.

What matters more is completeness. A dense feed full of holes won’t help much. Prioritize availability instead: 99%+ is strong, while anything below 95% is weak [1].

Error Risk

The biggest data-quality problems tend to show up in mixed fleets. A large NREL study covering 19,460 inverters found that only 41% passed data quality checks. The rest had problems such as data shifts (22%), missing data (5%), inconclusive orientation data (4%), and excessive clipping (2%) [1].

Vendor inconsistency makes this messier. Different manufacturers often use different field names for the same variable. For example, “Pac” and “ActivePower” can refer to the same thing. If your pipeline doesn’t account for that, it can misread the data. Firmware updates can cause trouble too. They may change which fields are available or shift scaling factors in the middle of deployment, which breaks historical time series.

This source becomes useful only after you normalize field names, units, and timestamps. The usual fix is a normalization layer that maps vendor-specific fields into one shared schema and tracks device metadata such as firmware version and communication protocol.

Those patterns get a lot more useful when you line them up with maintenance history.

Fit for Business Forecasting

Inverter logs and string-level data matter most for fault analysis, degradation tracking, and predictive maintenance. They’re not a standalone forecasting input. They work better as a diagnostic layer that explains why output drifts from expectations.

That matters because inverters account for 30% to 40% of all solar downtime, and telemetry can flag developing issues 7 to 14 days before total failure [1].

For degradation tracking, multi-year inverter and string datasets let you calculate performance ratios at the inverter or string level, then study those trends over time. That helps separate actual decline from normal seasonal swings. String-level trends can also reveal module aging or new shading long before plant-level data picks it up.

Use these signals to flag component-level loss, then confirm the root cause with asset records and maintenance history.

5. Asset Records and Maintenance History

Component-level fault data only helps forecasting when it’s tied back to the asset’s history. Asset records and maintenance history are static compared with telemetry, but they fill in things telemetry can’t show: how the plant was built, what changed over time, and how equipment has actually performed in service.

That means tracking equipment inventories, as-built string maps, commissioning reports, warranty documents, performance tests, and work orders. For each work order, capture the date, component, failure mode, root cause, parts replaced, downtime, and resolution. Without that context, service history is just a pile of records. With it, it becomes forecast input.

Coverage

Good coverage starts with full asset registration, a continuous work-order record, and logged outages or derates from commissioning through today. Include maintenance, repairs, updates, contracts, warranties, and external events such as grid outages and curtailment [18].

If those outside events are left out, production losses can get pinned on equipment issues by mistake. And once that happens, the forecast starts leaning in the wrong direction.

Latency and Resolution

Asset records move slowly on purpose. Static data, such as equipment models and nameplate capacity, usually changes only when a project is built, repowered, or reconfigured. So the lag is often measured in months or years.

Maintenance history updates more often, usually with each work order, often within a day or two after closure. Even so, it still reflects discrete events, not a continuous time series.

In practice, asset and maintenance data shape assumptions around degradation, availability, and repeat failures for medium- and long-range forecasts. SCADA and inverter logs deal with short-term swings. One handles the background story; the other handles the moment-by-moment signal. They work together, not against each other.

Error Risk

The biggest problems are inconsistent data entry and disconnected systems.

When technicians log work orders in different ways – skipping failure codes, using different names for the same part, or leaving downtime blank – the data gets messy fast. It becomes hard to group repeat issues or compare results across sites.

Disconnected systems make that worse. If asset registers, field service tickets, and financial records sit in separate tools that don’t talk to each other, analysts end up spending more time matching records than using them.

Use these records to separate actual production loss from planned work, grid events, and repowering changes. Common failure modes look like this:

Error TypeEffect on ForecastingFix
Missing or inconsistent failure codesCannot cluster recurring faults or identify high-risk asset classesEnforce controlled vocabularies in work-order templates
Unlogged minor outagesAvailability is overstated; forecasts are too optimisticIntegrate SCADA alarms with maintenance ticketing
Outdated equipment records after repoweringModels assume the wrong capacity or efficiency curvesTrigger asset record updates as part of the repowering workflow
No serial numbers on replaced componentsTraceability gaps for compliance and warrantiesRequire serial capture as a mandatory work-order field

Fit for Business Forecasting

Asset records and maintenance history support three forecast inputs that other sources can’t reliably supply: degradation rates, availability factors, and recurring failure risk.

Equipment specs and commissioning test results set the starting performance baseline. After that, long-term maintenance patterns – PID events, backsheet cracking, tracker wear – help shape site-level degradation rates instead of relying on generic assumptions. Historical downtime logs, broken out by failure class, then feed straight into annual availability calculations.

Blu Banyan’s SolarSuccess ERP puts asset registers, project records, work orders, inventory, and financials into one system [16][17]. When asset records are structured well, forecast assumptions are easier to trace and easier to audit. At that point, asset history stops being just admin paperwork and starts serving as part of the forecasting stack.

Pros and Cons of Each Data Source for Forecasting

No single source handles every forecasting job. Each one covers a different layer of the picture. Start with broad weather inputs, move to site sensors, then look at plant and lifecycle records. Remote weather data sets the baseline. Site sensors tune that baseline. Operational data helps explain where the gaps come from.

Data SourceMajor AdvantagesMajor DisadvantagesBest-Fit Forecasting Use Case
Satellite / Reanalysis Weather DataWide geographic coverage; long historical record; low-cost; no site hardwareLower spatial resolution; can miss local terrain effects; needs bias correction for site-level accuracyPortfolio screening, bankability studies, long-range production estimates, sites without instrumentation
On-Site Weather Stations & Irradiance SensorsCaptures site-specific conditions remote data cannot resolve; high-frequency samplingHigh upfront cost and maintenance burden; calibration drift; data gaps if poorly maintainedLender-grade yield validation, PPA compliance, local calibration of remote weather products
SCADA System DataHigh-frequency telemetry; strong for anomaly detectionDoes not isolate weather-driven variance; requires very high data completeness to support predictive modelsReal-time operations, performance ratio tracking, curtailment detection, forecast back-testing
Inverter Logs & String-Level DataGranular fault detection; early warning for inverter issues; captures string mismatchesVendor-specific formats; data can be noisy; limited value for broad weather-driven forecasting on its ownPredictive maintenance, root-cause analysis, degradation tracking
Asset Records & Maintenance HistorySupports degradation and availability modeling; explains long-term yield divergenceStatic and slow to update; often unstructured or incomplete; not a real-time inputLoss attribution, explanatory modeling, post-event performance analysis, long-term lifecycle planning

The best source depends on the job at hand. If the goal is planning, remote weather data and site sensors do most of the heavy lifting. If the goal is day-to-day plant oversight, SCADA and inverter data matter more. And if you’re trying to figure out why modeled yield and actual yield drift apart over months or years, asset records and maintenance history start to do the talking.

Put simply:

  • SCADA and inverter data push forecasting toward operations.
  • Asset records sit on the explanatory side of the stack. They move more slowly, but they matter when you need to understand long-term yield divergence.

Conclusion and Recommended Data Stack

No single source can do the whole job. The strongest yield models combine weather, sensor, operational, and asset data so teams can balance forecast accuracy, speed, and explainability. The right mix comes down to coverage, latency, error risk, and forecast horizon.

A better way to look at it is this: the stack isn’t a group of replacements. It’s a sequence. Start with the baseline, add calibration, bring in operations, then layer on context. For U.S. solar companies, the practical move is to use a tiered stack:

  • Tier 1 – Resource baseline: Satellite and reanalysis weather data
  • Tier 2 – Site calibration: On-site sensors for validation and bias correction
  • Tier 3 – Operational layer: SCADA and inverter logs for production tracking and anomaly detection
  • Tier 4 – Context layer: Asset records and maintenance history to explain long-term divergence

That setup lines up well with the most common use cases:

Business GoalRecommended Data SourcesWhy It Fits
Long-term yield estimation and bankabilitySatellite/reanalysis + on-site irradiance sensorsBest mix of history and site truth
Performance guarantees and contractual reportingOn-site sensors + SCADA availability metrics + asset recordsDefensible irradiance and normalized production data
Short-term operational forecastingSCADA + inverter logs + on-site sensorsHigh-frequency telemetry shows current plant status
Predictive maintenance and fault detectionSCADA + inverter/string-level logs + maintenance historyHelps isolate failure causes
Portfolio risk managementSatellite/reanalysis + SCADA + asset recordsBroad coverage helps flag higher-variability or higher-risk sites
Financial reporting and revenue planningSCADA production data + asset records + ERP financial recordsConnects production data to finance

When these data sets sit in different systems, achieving project management excellence through system integration becomes the last piece of the puzzle. Blu Banyan’s SolarSuccess can connect SCADA, work orders, inventory, and financial records in one system, while keeping asset and maintenance history structured for forecasting and reporting.

FAQs

Which data source should I prioritize first?

Start with solar irradiance, insolation, and solar spectrum data in your ERP to set a reliable performance baseline. That gives you a clean way to tell the difference between output swings caused by weather and drops caused by equipment trouble.

For operations and predictive maintenance, put inverter data first. Inverters account for 30% to 40% of solar downtime, and they generate detailed fault logs that can point to trouble early. After that, bring in device telemetry, revenue-grade meters, and asset master data.

How often should solar yield model inputs be updated?

It depends on the data type and the role it plays in day-to-day operations.

For live monitoring and grid management, SCADA and sensor data are usually logged every 1 to 5 minutes. In high-availability setups, polling often happens every 5 to 15 seconds.

Financial and project inputs should update automatically as milestones happen. And if you’re building baseline performance models, you’ll usually need 6 to 12 months of clean historical data to work from.

For reporting, daily or monthly updates should be automated so teams have steady, consistent visibility.

What causes the biggest errors in solar yield models?

The biggest errors usually start with inconsistent or low-quality data inputs. And that’s a problem, because bad data can mask actual equipment issues.

Some of the most common causes are telemetry loss, faulty sensor signals such as incorrect Modbus mapping or negative irradiance readings, disconnected systems that prevent accurate reconciliation, and mismatched data formats for dates, units, or currency.

Without automated validation, these problems can ripple across dashboards. Once that happens, it gets much harder to tell whether a drop in performance came from the weather or from a technical fault.

Illustration: Community with energy efficient buildings, solar panel array, wind turbines, trees, flowers, and people riding bicycles.