Fecha de publicación:
--
Fuente:
PubMed "swarm"
Biomimetics (Basel). 2026 Sep 21;11(9):682. doi: 10.3390/biomimetics11090682.ABSTRACTLearning-based hyper-heuristics for combinatorial optimization select from among low-level operators or tune the parameters of a single metaheuristic. However, online selection among complete population-based metaheuristics that share one evolving population and its joint configuration through a common action space have not been formulated or evaluated for binary combinatorial optimization problems. This work presents such a formulation and evaluates it on two binary domains: the multidimensional knapsack problem (MKP), and the set covering problem (SCP). A proximal policy optimization (PPO) agent orchestrates a portfolio of seven nature-inspired population-based metaheuristics, transferring the full population between methods at each decision, then acts at every fixed quantum of iterations on a fourteen-dimensional size-independent description of the search state. Over the same state, reward, and portfolio, three action spaces of increasing richness are instantiated: discrete selection, joint selection and configuration through algorithm-specific translation maps, and heterogeneous co-evolution through a simplex allocation. The three techniques are evaluated under five-fold cross-validation with twenty recorded seeds on the Chu and Beasley cb9 set (n=500 items, B=25,000 evaluations, quantum Q=25) against the best-known values of the recent literature, and the selection technique is also evaluated on 65 set-covering instances. On cb9, the selection agent is ahead of the blind round-robin and random baselines. The co-evolution agent places second, with a 0.043 percentage point gap to the best-known value behind the strongest single metaheuristic. At half the evaluation budget, the two agents rank first and second; on the Set Covering Problem, where no single method dominates, the selection agent ranks first among eleven strategies. An untrained control and a portfolio ablation show that the policy learns mainly which members to avoid; the joint selection-and-configuration agent does not yet improve on the blind baselines. The formulation, translation maps, simplex allocation, and evaluation protocol are released with this study.PMID:42782708 | DOI:10.3390/biomimetics11090682