Are you interested in building the data infrastructure that modern products and AI systems depend on?
About the product
Many of the systems Forma Pro works with depend on data coming from multiple APIs, databases, SaaS platforms, event streams, documents and legacy systems.
Before analytics, machine learning or AI can work reliably, this data has to be collected, normalized, validated and made available with predictable quality and latency.
Our projects include operational data platforms, integration layers, event-driven systems and data pipelines for marketplaces, logistics, MarTech, PropTech and other complex digital businesses.
Technologies that we use
- Python
- SQL
- PostgreSQL
- Kafka and event-driven architectures
- Airflow and workflow orchestration
- dbt
- AWS
- Docker
- APIs, webhooks and third-party integrations
- Analytical and operational data stores
Individual projects may use different infrastructure depending on their requirements.
About you
- You enjoy making messy data understandable and dependable
- You treat reliability, testing and observability as part of data engineering rather than optional extras
- You think carefully about schemas, contracts and how systems fail
- You are comfortable integrating third-party platforms that were clearly not designed with your convenience in mind
- You can work independently and take responsibility for architecture decisions
- You prefer sustainable systems over heroic manual fixes
Required skills and competencies
- Spoken English — upper-intermediate or above
- 4+ years of professional data engineering or backend engineering experience
- Strong Python and SQL skills
- Experience designing and maintaining production data pipelines
- Experience working with relational databases and analytical data stores
- Experience integrating APIs, webhooks and external data sources
- Understanding of batch and streaming architectures
- Understanding of data modelling, schema design and data quality
- Experience with cloud infrastructure
- Automated testing experience
Would be a plus
- Kafka or another event-streaming platform
- Airflow, Dagster, Prefect or similar orchestration tooling
- dbt
- Snowflake, BigQuery, Redshift or Databricks
- IoT or high-frequency event ingestion
- Data observability and lineage
- Reverse ETL
- Vector databases and embedding pipelines
- Experience preparing data infrastructure for ML or LLM applications
Responsibilities
- Design and build scalable data ingestion and transformation pipelines
- Integrate external APIs, databases, SaaS systems, event streams and files into reliable data flows
- Build batch, streaming and near-real-time processing systems
- Define data models, schemas and contracts between services
- Create normalization layers for heterogeneous data sources
- Implement automated data validation and quality checks
- Build monitoring for freshness, schema drift, failed jobs and unusual data behaviour
- Improve performance, reliability and maintainability of existing pipelines
- Work with AI and ML engineers to provide dependable datasets for models and agents
- Participate in architecture decisions and help choose appropriate infrastructure for each project
- Own data systems from initial design through production operation
What we offer
- No micromanagement and no working-time surveillance
- Freedom to make engineering decisions and choose appropriate tools
- Flexible working hours
- Fully remote work
- Complex integration and data problems rather than endless CRUD tasks
- An opportunity to work closely with AI, ML and backend engineers
- Direct influence on architecture
- Company support when you need it
- Direct communication with Forma Pro's CEO and CTO