Overview
Oddball is looking for a Data Engineer to join its SEC team, focusing on building and maintaining data pipelines and infrastructure for AI-driven information discovery within a federal financial regulatory agency.
Key Responsibilities
- Build and maintain data pipelines for NLQ, semantic search, and knowledge graph capabilities on the SEC's Enterprise Data Platform.
- Onboard new data sources and extend the semantic layer without disrupting existing pipelines.
- Ingest, transform, and deliver XBRL and SEC financial filing datasets.
- Support GraphRAG and knowledge graph pipelines alongside the Data Architect.
- Work with AWS data services, including Athena, DataZone, SageMaker, S3, and Lambda.
- Implement and maintain vector database and semantic search tooling.
- Manage metadata, data lineage, and data quality validation across the pipeline ecosystem.
- Monitor pipeline health and resolve data integration issues across environments.
- Ensure compliance with federal security and requirements.
Requirements
- Hands-on experience building enterprise data pipelines in regulated or federal environments.
- Strong SQL proficiency and experience with PostgreSQL or similar databases.
- Experience with ETL/ELT tools such as AWS Glue or Airflow.
- Proficiency in Python for pipeline development and automation.
- Familiarity with vector databases or semantic search tooling.
- Experience with graph databases such as Neptune or Neo4j (a plus).
- Understanding of FISMA and NIST 800-53 compliance frameworks.
- Clear communication skills for documenting pipeline logic and collaborating with the Data Architect.
Benefits
- Fully remote work.
- Tech & Education Stipend.
- Comprehensive Benefits Package.
- Company Match 401(k) plan.
- Flexible PTO and Paid Holidays.
Location
Remote within the United States.
How to Apply
To apply for this position, please visit the company's career page or send your application to hello@Oddball.io.
Deadline
No specific deadline mentioned.