Advanced Job Skill Forecasting with Job-SDF: A New Era in Workforce Planning
Researchers from leading Chinese institutions have developed Job-SDF, a comprehensive dataset derived from 10.35 million job ads, to improve job skill demand forecasting using advanced machine learning models. This dataset aims to enhance the alignment of workforce skills with market needs, supporting economic growth and stability.
In the dynamic and rapidly evolving job market, the ability to forecast skill demand is critical for policymakers and businesses to stay ahead of changes, ensuring that workforce skills are aligned with market needs. This alignment enhances productivity and competitiveness and directs individuals toward relevant training and education opportunities, fostering continuous self-learning and development. However, the significant challenge in advancing this field lies in the absence of comprehensive datasets. To bridge this gap, researchers from the University of Science and Technology of China, Career Science Lab at BOSS Zhipin, PBC School of Finance at Tsinghua University, Institute of Artificial Intelligence at Beihang University, and AIT at HKUST (GZ) have introduced Job-SDF, a multi-granularity dataset designed specifically for training and benchmarking job-skill demand forecasting models.
Introducing Job-SDF: A Comprehensive Dataset for Skill Forecasting
Job-SDF is derived from 10.35 million public job advertisements collected from major online recruitment platforms in China between 2021 and 2023. The dataset encompasses monthly recruitment demand for 2,324 types of skills across 521 companies, uniquely enabling the evaluation of skill demand forecasting models at various levels, including occupation, company, and regional granularities. This extensive dataset and the accompanying benchmarking tools are publicly accessible via a dedicated GitHub repository, providing a valuable resource for researchers and practitioners in the field.
From Surveys to Online Platforms: Transforming Skill Demand Analysis
Traditional skill demand analysis has relied heavily on labor-intensive, survey-based methods, often limited to specific companies or occupations. The rise of online recruitment platforms has transformed this landscape, accumulating vast amounts of job advertisement data. By leveraging this data, researchers have formulated skill demand forecasting as a time series task, utilizing various machine learning models such as autoregressive integrated moving average (ARIMA), recurrent neural networks (RNNs), and dynamic graph autoencoders (DyGAEs) to predict future skill needs. However, the lack of comprehensive and publicly accessible datasets has been a significant impediment, making it difficult to replicate experimental results and identify bottlenecks in current research. Existing datasets typically focus on predicting skill demand variations across different occupations, with a notable lack of modeling and prediction at other granularities, such as companies or regions. This limitation hinders comprehensive comparisons between different models and impedes the exploration of potential downstream applications, such as human capital strategy development and regional policy formulation.
Building the Dataset: A Rigorous Process
To construct the Job-SDF dataset, researchers collected job advertisements, extracted skill terms using a Named Entity Recognition (NER) model, and quantified monthly skill demand at various granularities. The data collection process involved annotating a dataset to train the NER model by identifying skill terms within job requirement texts. This annotation was followed by a series of steps to create a refined skill dictionary, which was then used to filter and map skill words extracted by the NER model, resulting in standardized skill requirements for each job advertisement. The demand for different skills in the job market was estimated based on the volume of job advertisements listing these specific skills as requirements within a given time period.
Benchmarking Models: Insights and Performance
Benchmarking various models on the Job-SDF dataset provided valuable insights into their performance. Traditional statistical models like ARIMA and Prophet showed limitations in capturing complex nonlinear relationships and exhibited suboptimal performance in large-scale data scenarios. RNN-based models, particularly those employing segment-wise iterations like SegRNN, demonstrated enhanced performance by reducing the recurrence count within RNNs. Transformer-based models, despite their global modeling capabilities, struggled with the shorter time series context. However, models like PatchTST, which focus on local information, showed significant improvements. The performance of different linear models varied, with DLinear outperforming most Transformer-based models, while TSMixer performed poorly, possibly due to overfitting. Graph-based models like CHGH and Pre-DyGAE exhibited poor performance in separate skill demand forecasting scenarios, likely due to a mismatch between their model design and the context of the dataset. FiLM, which employs a denoising-based model, achieved the best performance in most cases, demonstrating robustness.
Further analysis of the dataset revealed the varying nature of skill demand values and the impact of structural breaks in job skill demand time series data. Structural breaks, which indicate significant changes in the statistical properties of skill demand over time, pose a challenge for forecasting models. The Chow test was used to detect these breaks, and models like FiLM demonstrated robustness in mitigating their disruptive impact. The analysis also showed that while forecasting performance on skills experiencing structural breaks was generally worse, models like FiLM maintained stable and reliable predictions across various granularities and demand value ranges.
The Job-SDF dataset provides a robust framework for advancing research in job skill demand forecasting. By offering a comprehensive dataset and benchmarking tools, it facilitates the development of more accurate and adaptable forecasting models. This, in turn, supports better alignment of workforce skills with market needs, promoting economic growth and stability. Job-SDF addresses a critical gap in the field, enabling researchers to replicate results, conduct comprehensive comparisons, and explore new applications in human capital strategy and regional policy development. As a publicly accessible resource, it holds significant potential to drive forward research and practical applications in job skill demand forecasting.
- FIRST PUBLISHED IN:
- Devdiscourse
Google News