نوع مقاله : مقاله پژوهشی

نویسندگان

1 گروه مهندسی نقشه برداری، دانشکده مهندسی عمران آب و محیط‌زیست، دانشگاه شهید بهشتی، تهران، ایران

2 گروه مهندسی نقشه برداری، دانشکده مهندسی عمران، آب و محیط‌زیست، دانشگاه شهید بهشتی، تهران، ایران

چکیده

پیشینه و اهداف: باتوجه‌به افزایش جمعیت، فشار بر منابع طبیعی، تغییرات اقلیمی و اهمیت روزافزون امنیت غذایی، پیش‌بینی دقیق عملکرد محصولات کشاورزی، به‌ویژه گندم، به یکی از موضوعات مهم پژوهشی تبدیل شده است. در سال‌های اخیر، استفاده از داده‌های سنجش‌ازدور، اطلاعات اقلیمی و ویژگی‌های خاک در کنار الگوریتم‌های یادگیری ماشین، فرصت‌های جدیدی برای برآورد عملکرد محصول فراهم کرده است. بااین‌حال، بسیاری از پژوهش‌های پیشین یا تنها بر یک منبع داده تمرکز داشته‌اند یا در مقیاس‌های محدود مکانی و زمانی انجام شده‌اند؛ بنابراین هنوز شکافی در زمینه ارزیابی جامع و مقایسه‌ای مدل‌ها با استفاده از داده‌های چندمنبعی وجود دارد. هدف این پژوهش، بررسی توانایی الگوریتم‌های مختلف یادگیری ماشین در پیش‌بینی عملکرد گندم با بهره‌گیری از داده‌های اقلیمی، شاخص NDVI حاصل از سنجش‌ازدور و همچنین ارزیابی قابلیت تعمیم این مدل‌ها در شرایط مکانی و اقلیمی متفاوت بوده است.
روش‌ها‌: این پژوهش بر اساس داده‌های چندمنبعی شامل متغیرهای اقلیمی و شاخص‌ سنجش‌ازدور NDVI انجام شد. مطالعه در ۹ منطقه مختلف در آلمان و در بازه‌ای ۲۲ساله صورت‌گرفته و بدین ترتیب امکان بررسی تغییرات مکانی و زمانی عملکرد گندم فراهم شده است. در این تحقیق از الگوریتم‌های XGBoost، Random Forest و Ridge Regression برای پیش‌بینی عملکرد گندم استفاده شده است. از میان متغیرهای مؤثر، شاخص NDVI و تنش تبخیری از جمله عوامل مهم در تحلیل مدل‌ها بوده‌اند. عملکرد مدل‌ها از نظر دقت پیش‌بینی و میزان خطا مورد ارزیابی قرار گرفته و شاخص‌هایی مانند ضریب تعیین، ریشه میانگین مربعات خطا و میانگین خطای مطلق برای مقایسه نتایج استفاده شده‌اند. طراحی مطالعه به‌گونه‌ای بوده که علاوه بر تحلیل داده‌ها، امکان مقایسه عملکرد مدل‌های کلاسیک و پیشرفته در سطح منطقه‌ای نیز فراهم شود.
یافته‌ها: نتایج پژوهش نشان داد که الگوریتم XGBoost نسبت به دو مدل Random Forest و Ridge Regression عملکرد بهتری در پیش‌بینی عملکرد گندم داشته است. این مدل توانست الگوهای فضایی و زمانی عملکرد را با خطای کمتر و قدرت تبیین بالاتر بازسازی کند. همچنین مشخص شد که متغیرهای NDVI، تنش تبخیری نقش مهمی در شکل‌گیری الگوهای عملکرد دارند. در مناطق یکنواخت‌تر و پربازده‌تر، دقت مدل بالاتر بود، درحالی‌که در مناطقی با ناهمگنی محیطی بیشتر، مانند برخی نواحی خاص، خطای پیش‌بینی افزایش یافت. باوجوداین، روند کلی نشان داد که روش‌های یادگیری ماشین، به‌ویژه مدل‌های مبتنی بر تقویت گرادیان، توان بالایی در شناسایی روابط پیچیده بین عوامل محیطی و عملکرد محصول دارند.
نتیجه‌گیری: در مجموع، نتایج این پژوهش نشان می‌دهد که ترکیب داده‌های سنجش‌ازدور و اقلیمی همراه با الگوریتم‌های پیشرفته یادگیری ماشین می‌تواند ابزاری مؤثر برای پیش‌بینی عملکرد گندم و پشتیبانی از تصمیم‌گیری در کشاورزی هوشمند باشد. برتری XGBoost نسبت به سایر مدل‌ها حاکی از آن است که مدل‌های مبتنی بر یادگیری ماشین توانایی بالایی در تحلیل روابط غیرخطی و چندبعدی موجود در داده‌های کشاورزی دارند. بااین‌حال، محدودیت‌هایی نیز وجود دارد؛ از جمله نبود داده‌های مدیریتی مزرعه مانند تاریخ کاشت، نوع رقم، میزان آبیاری و مصرف کود، و نبود داده‌های خاکی با تفکیک مکانی و زمانی کافی برای ورود به مدل نهایی، از محدودیت‌های پژوهش حاضر بودند. همچنین استفاده از مدل‌های پیچیده‌تر ممکن است در شرایط داده محدود با خطر بیش‌برازش همراه باشد؛ بنابراین، پیشنهاد می‌شود در پژوهش‌های آینده از مجموعه‌داده‌های کامل‌تر، مدل‌های هیبریدی فضایی - زمانی و روش‌های پیشرفته‌تر مانند شبکه‌های عصبی بازگشتی و کانولوشنی استفاده شود. این رویکرد می‌تواند به افزایش دقت پیش‌بینی، تعمیم‌پذیری بیشتر و کاربرد عملی بهتر در مدیریت پایدار کشاورزی و امنیت غذایی کمک کند.

کلیدواژه‌ها

موضوعات

عنوان مقاله [English]

Yield prediction of winter wheat using remote sensing data and machine‑learning algorithms

نویسندگان [English]

  • F. Babayi 1
  • A. Vafaeinejad 2
  • A. Sharifi 1

1 Department of Surveying and Geoinformatics Engineering, Faculty of Civil, Water and Environmental Engineering, Shahid Beheshti University, Tehran, Iran

2 Department of Surveying and Geoinformatics Engineering, Faculty of Civil, Water and Environmental Engineering, Shahid Beheshti University, Tehran, Iran

چکیده [English]

Background and Objectives: With the growing global population, increasing pressure on natural resources, climate change, and the rising importance of food security, accurate crop yield prediction—particularly for wheat—has become a crucial research topic. In recent years, the integration of remote sensing data, climatic information, and soil properties with machine‑learning algorithms has created new opportunities for estimating crop yields. However, many previous studies have either focused on a single data source or have been conducted at limited spatial and temporal scales. Consequently, a gap still exists in conducting comprehensive and comparative evaluations of models using multi‑source datasets. The aim of this study was to evaluate the capability of different machine-learning algorithms to predict wheat yield using climatic variables and the remotely sensed Normalized Difference Vegetation Index (NDVI), and to assess the generalizability of these models across different spatial and climatic conditions.
Methods: This study was conducted using multi-source data comprising climatic variables and the remotely sensed Normalized Difference Vegetation Index (NDVI). The study covered nine regions in Germany over a 22-year period, thereby enabling the assessment of spatial and temporal variations in wheat yield. The XGBoost, Random Forest, and Ridge Regression algorithms were employed to predict wheat yield. Among the investigated variables, NDVI and vapor pressure deficit (VPD), as an indicator of atmospheric evaporative stress, were identified as important predictors in the model analyses. Model performance was evaluated in terms of predictive accuracy and error using the coefficient of determination (R²), root mean square error (RMSE), and mean absolute error (MAE). The study was designed not only to analyze the data but also to enable a regional comparison between conventional and advanced predictive models.
Findings: The results showed that the XGBoost algorithm outperformed both Random Forest and Ridge Regression in predicting wheat yield. XGBoost was able to reconstruct spatial and temporal yield patterns with lower error and higher explanatory power. The findings also indicated that NDVI and vapor pressure deficit (VPD), as an indicator of atmospheric evaporative stress, were important predictors of wheat yield patterns. Higher model accuracy was observed in more homogeneous and high‑yielding regions, whereas prediction errors increased in areas with greater environmental heterogeneity. Nevertheless, the overall trend demonstrated that machine‑learning techniques—especially gradient‑boosting models—have strong capabilities in capturing the complex relationships between environmental factors and crop yield.
Conclusion: Overall, the findings of this study demonstrate that integrating remote-sensing and climatic data with advanced machine-learning algorithms can provide an effective tool for predicting wheat yield and supporting decision-making in smart agriculture. The superior performance of XGBoost compared with the other models indicates its strong capability to capture the nonlinear and multidimensional relationships inherent in agricultural data. Nevertheless, this study had several limitations, including the lack of farm-management data, such as sowing dates, cultivar types, irrigation levels, and fertilizer application rates, as well as the absence of soil data with sufficient spatial and temporal resolution for inclusion in the final models. Moreover, the use of more complex models may increase the risk of overfitting when data are limited. Therefore, future studies are recommended to employ more comprehensive datasets, hybrid spatiotemporal models, and advanced approaches such as recurrent neural networks and convolutional neural networks. Such approaches could improve predictive accuracy and model generalizability while enhancing the practical application of yield-prediction models in sustainable agricultural management and food-security planning.

کلیدواژه‌ها [English]

  • Winter Wheat Yield
  • Remote Sensing
  • XGBoost
  • SHAP
  • NDVI
  • VPD

COPYRIGHTS

© 2026 The Author(s).  This is an open-access article distributed under the terms and conditions of the Creative Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)

(https://creativecommons.org/licenses/by-nc/4.0/)

doi: 10.1016/j.heliyon.2024.e40836
doi: 10.32604/cmc.2024.050240
doi: 10.3390/agronomy16060670
doi: 10.1080/01431161.2021.1974116
doi: 10.48550/arXiv.2306.04566
doi: 10.1016/j.procs.2023.01.023
doi: 10.1080/01140671.2022.2032213
doi: 10.1109/JSTARS.2024.3435699
doi: 10.1016/j.compag.2023.107663
doi: 10.22067/jcesc.2025.93323.1395
doi: 10.1016/j.jag.2024.104183
doi: 10.1016/j.agwat.2026.110254
doi: 10.1038/s41598-023-30813-7
[21] Liaghat A, Dehghani T, Rezaei Rad H, Ahmadpari H. [Estimating grain corn yield based on Landsat 8 satellite images (case study: Shahid Beheshti Agro-Industry lands of Dezful)]. Iranian Journal of Irrigation and Drainage. 2024; 18(3): 433–447. [In Persian]
doi: 10.22125/agmj.2021.266547.1107
[24] Amir Ashayeri E, Rezaverdinejad V, Bahmanesh J, Asadzadeh F, Rahimi M. [Modeling and forecasting rainfed wheat yield based on meteorological variables using combined artificial intelligence methods]. Iranian Journal of Irrigation and Drainage. 2025; 19(2): 243–261. [In Persian]
doi: 10.22077/jdcr.2024.7942.1072
doi: 10.22069/ejcp.2019.14029.2066
doi: 10.3389/fpls.2024.1451607
doi: 10.48550/arXiv.2512.15140
doi: 10.3390/w14081188
doi: 10.1016/j.eja.2018.09.003
doi: 10.1109/ACCAI61061.2024.10601854
doi: 10.3390/su16188277