Ugrás a tartalomra
Vissza a hírekhez
Hugging Face2026. szept. 28.eszköz

Közös gyűjtőhelyet kaptak az AI-ágenseket tanító környezetek

A Hugging Face Hub új felületet indított a megerősítéses tanulási (RL) környezeteknek, így egyszerűbbé válik az AI-ágensek fejlesztése és tesztelése.

Kipróbálom: Hugging Face RL Environmentshuggingface.co/datasets?other=rl-environment
Welcome RL Environments to the hub

A Hugging Face bejelentette, hogy mostantól dedikált helyet biztosít a megerősítéses tanulási (RL) környezeteknek a platformján. Az RL Environments szekcióban a fejlesztők 1 közös helyen érhetik el a különböző keretrendszerekhez készült feladatokat és teszteket. Ezzel megszűnik a korábbi töredezettség, ahol minden fejlesztőcsapat saját egyedi adatbázist használt.

A megerősítéses tanulás során az AI-ágensek feladatokat kapnak, amelyekre adott válaszaikért pontszámot, azaz jutalmat kapnak. A Hugging Face megoldása lehetővé teszi, hogy a fejlesztők 1 kattintással futtassák ezeket a környezeteket olyan népszerű keretrendszerekben, mint a Harbor, a Verifiers vagy az NVIDIA NeMo Gym. A hibajavítások is egyszerűbbé válnak, hiszen a javítások minden keretrendszerre azonnal érvényesülnek.

Az új funkció használatához a fejlesztőknek mindössze egy egyszerű címkét kell elhelyezniük az adatbázisuk leírásában. Az így megjelölt környezetek azonnal megjelennek a nyilvános, böngészhető listában, megkönnyítve az AI-ágensek teljesítményének mérését és összehasonlítását.

Az eredeti szöveg (Hugging Face)
Stop building environment registries What shipped Run an environment and inspect its reward Harbor: run a reference solution Verifiers: run a model on the same task OpenEnv: run an agent and inspect its reward NeMo Gym: generate responses and inspect rewards Tag your environment Already on the Hub What comes next Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments. An environment gives an agent a task, responds to its actions with observations, and scores the outcome. The resulting rewards can measure an agent's performance during evaluation or provide a learning signal during training. For an introduction to this interaction loop, see our blogpost on environments. Within the environment, the agent will perform a set of tasks that are represented as datasets. Therefore, environments can be split into broadly two parts: tasksets and runtimes. In this release, we are focusing on the tasksets. An RL environment on the Hub is a dataset repo that shows up in the new RL Environments filter. The Use this dataset button gives you the command to run it in that framework. There is no new repo type, no registry, and no sign-up. There are already environments in Harbor, Verifiers, and NVIDIA NeMo Gym. Every RL paper or framework uses its own way to find environments. Custom hubs, runtime registries, independent task datasets, or a GitHub list of tasks with a custom loader. This means that many of the published environments are siloed: if you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand. We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say "this data is an environment, and here is how you run it." The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub. The dataset repository hosts your environment files. The framework runs them locally or on a supported cloud backend. Hugging Face Jobs can run cloud workloads, and Hugging Face Sandboxes, built on Jobs, provide interactive command execution. The tags describe compatibility and generate loading commands; adding a tag does not start a job or sandbox. The RL Environments filter. Go to huggingface.co/datasets?other=rl-environment. Every dataset with the rl-environment tag appears there, whatever framework it works with. Framework tags. Four environment frameworks are registered as dataset libraries: Each framework tag puts the framework's icon on the dataset page and adds a generated snippet to Use this dataset. A dataset can carry more than one framework tag. That is the point. Tags describe compatibility, and compatibility is not exclusive. Each listed framework must support the files in the repository; adding a tag does not convert them. Choose the example for your framework and run it in a separate Python environment with the prerequisites listed below. Harbor can load task directories from a Hub repository. The oracle agent runs the task's reference solution, then the verifier scores the result. It does not call a model. The viewer shows the task's reward, verifier output, and logs. This checks the task and its reference solution before you try a model agent. The Harbor integration of verifiers v1 can run the same task directories in different runtimes, such as Docker. It also supports different harnesses, including a minimal bash