پیش‌بینی مقاومت گیاهچه‌های برنج به سمیت آهن با استفاده از هوش مصنوعی و داده‌های مولکولی و فنوتیپی

نوع مقاله : مقاله پژوهشی

نویسندگان

1 دانشجو کارشناسی ارشد، گروه تولیدات گیاهی، دانشکده علوم کشاورزی و منابع طبیعی، دانشگاه گنبدکاووس، ایران.

2 استادیار، گروه تولیدات گیاهی، دانشکده علوم کشاورزی و منابع طبیعی، دانشگاه گنبدکاووس، ایران.

3 استاد، گروه تولیدات گیاهی، دانشکده علوم کشاورزی و منابع طبیعی، دانشگاه گنبدکاووس، ایران.

چکیده

هدف: سمیت آهن یکی از تنش‌های مهم محدودکننده تولید برنج در خاک‌های غرقابی اسیدی است. روش‌های مرسوم ارزیابی تحمل به این تنش که مبتنی بر فنوتیپ‌برداری پرهزینه و زمان‌بر هستند، به‌عنوان یک محدودیت اساسی در برنامه‌های اصلاحی محسوب می‌شوند. این مطالعه با هدف ارائه یک چارچوب پیش‌بینی کارآمد و بین شرایطی انجام شد تا امکان ارزیابی تحمل به سمیت آهن را با استفاده از داده‌های جمع‌آوری‌شده در شرایط نرمال و تنش فراهم کند.
مواد و روش‌ها: برای این منظور، داده‌های فنوتیپی و ژنوتیپی (نشانگرهای iPBS، IRAP، ISSR و SSR) مربوط به ۹۶ لاین اینبرد برنج و والدین آن‌ها، در هر دو شرایط نرمال و تنش سمیت آهن گردآوری شد. پس از پیش‌پردازش داده‌ها شامل حذف داده گمشده، تشخیص و حذف نقاط پرت و نرمال‌سازی داده‌ها، انتخاب ویژگی با سه الگوریتم شامل آزمون F، الگوریتم MRMR و الگوریتم RReliefF برای وزن‌دهی ویژگی‌ها بر اساس توانایی آن‌ها در تمایز بین نمونه‌های همسایه مورد استفاده قرار گرفت. عملکرد ده مدل یادگیری ماشین شامل شبکه عصبی مصنوعی، ماشین بردار پشتیبان، درخت تصمیم، روش‌های گروهی شامل جنگل تصادفی، XGBoost و AdaBoost، فرایند گاوسی، رگرسیون خطی، تحلیل ممیز خطی نزدیک‌ترین همسایه و Naive Bayes در دوازده حالت پیش‌بینی مدل‌سازی مختلف (مبتنی بر نوع داده و شرایط آموزش - آزمون) با معیارهای MAE، RMSE و R² ارزیابی شد. داده‌ها به نسبت ۷۰:۳۰ به مجموعه‌های آموزش و آزمون تقسیم شدند و به‌منظور ارزیابی پایداری مدل‌ها و جلوگیری از بیش‌برازش، از اعتبارسنجی متقاطع استفاده شد. بهینه‌سازی فراپارامترها با استفاده از جستجوی شبکه‌ای و بهینه‌سازی بیزی انجام شد و ارزیابی نهایی مدل‌ها بر روی مجموعه آزمون صورت گرفت.
نتایج: نتایج نشان داد که حالت‌های پیش‌بینی مبتنی بر ادغام کامل داده‌ها (D10 و D12) با دستیابی به ضرایب تعیین بالا (R² > 0.96) به‌عنوان دقیق‌ترین رویکردها شناخته شدند. حالت پیش‌بینی D8 که تنها از داده‌های ژنوتیپی شرایط نرمال و الگوریتم MRMR استفاده می‌کند، با دقت متوسط (R² = 0.678) نشان داد که می‌توان برخی الگوهای آماری نشانگری مرتبط با پاسخ ژنوتیپ‌ها به سمیت آهن را در جمعیت مورد مطالعه شناسایی کرد و از آن‌ها برای غربالگری اولیه بین‌شرایطی استفاده نمود.
نتیجه‌گیری: نتایج این پژوهش نشان می‌دهد که استفاده مرحله‌ای از مدل‌ها می‌تواند در غربالگری اولیه ژنوتیپ‌ها مؤثر باشد؛ به‌طوری‌که مدل‌هایی مانند D8 با وجود دقت متوسط، می‌توانند به کاهش هزینه و تمرکز منابع در مراحل ابتدایی کمک کنند، در‌حالی‌که انتخاب نهایی ژنوتیپ‌های متحمل همچنان نیازمند اعتبارسنجی با مدل‌های دقیق‌تر، فنوتیپ‌برداری مستقیم تحت تنش و آزمون‌های مستقل مزرعه‌ای است.

کلیدواژه‌ها


عنوان مقاله [English]

Prediction of rice seedlings’ tolerance to iron toxicity using artificial intelligence, and molecular and phenotypic data

نویسندگان [English]

  • Milad Aliyani Nejad 1
  • Sayed Javad Sajadi 2
  • Hossein Sabouri 3
  • Mehdi Zarei 2
1 MSc Student, Department of Crop Production, Faculty of Agriculture Sciences and Natural Resources, Gonbad Kavous University, Iran.
2 Assistant Professor, Department of Crop Production, Faculty of Agriculture Sciences and Natural Resources, Gonbad Kavous University, Iran.
3 Professor, Department of Crop Production, Faculty of Agriculture Sciences and Natural Resources, Gonbad Kavous University, Iran.
چکیده [English]

Objective
Iron toxicity is a major constraint to rice production in acidic flooded soils. Conventional methods for evaluating tolerance to this stress rely on costly and time-consuming phenotyping, which limits their efficiency in breeding programs. This study aimed to develop an efficient cross-condition prediction framework to assess rice tolerance to iron toxicity using phenotypic and genotypic data collected under both normal and stress conditions.
Materials and Methods
Phenotypic and genotypic data derived from iPBS, IRAP, ISSR, and SSR markers were collected from 96 rice inbred lines and their parents under both normal and iron-toxicity stress conditions. Following data preprocessing, including removal of missing values, outlier detection, and normalization, feature selection was performed using three algorithms: F-test, MRMR, and RReliefF. Eleven machine-learning models, including artificial neural network (ANN), support vector machine (SVM), decision tree (DT), random forest (RF), XGBoost, AdaBoost, Gaussian process (GP), linear regression (LR), linear discriminant analysis (LDA), k-nearest neighbors (KNN), and Naive Bayes (NB), were evaluated across twelve prediction scenarios defined by data type and training–testing conditions. Model performance was assessed using MAE, RMSE, and R2R^2R2. The dataset was split into training and test sets at a 70:30 ratio, cross-validation was applied to improve model robustness, and hyperparameters were optimized using grid search and Bayesian optimization.
Results
Prediction scenarios based on full data integration (D10 and D12) showed the highest predictive accuracy, with coefficients of determination exceeding 0.96. Scenario D8, which relied solely on genotypic data from normal conditions and MRMR-based feature selection, achieved moderate predictive accuracy (R2=0.678R^2 = 0.678R2=0.678). This result suggests that marker-derived patterns associated with genotypes’ responses to iron toxicity can be partially captured from normal-condition data and used for preliminary cross-condition screening.
Conclusion
The results indicate that a stepwise prediction strategy can improve the efficiency of preliminary genotype screening for iron-toxicity tolerance. Models such as D8, despite their moderate accuracy, may help reduce costs and optimize resource allocation during early screening stages. However, final selection of tolerant genotypes still requires validation using more accurate models, direct phenotyping under stress conditions, and independent field trials.

کلیدواژه‌ها [English]

  • Rice
  • Iron toxicity
  • Machine learning
  • Between-condition prediction
  • Feature selection