mirror of
https://github.com/ada-dmitry/CourseWork_IRFM.git
synced 2026-09-24 01:10:18 +00:00
Add web crawler for legal documents
- Implemented a new crawler module to extract text and links from legal documents. - Added a loader module to handle HTML downloads with error handling and retries. - Created a main script to initiate the crawling process on a specified URL. - Defined functions for normalizing URLs, extracting links, checking terminal pages, and extracting text content. - Updated project metadata with dependencies and versioning in pyproject.toml and uv.lock.
This commit is contained in:
@@ -0,0 +1,9 @@
|
||||
[project]
|
||||
name = "coursework-irfm"
|
||||
version = "0.1.0"
|
||||
description = "Add your description here"
|
||||
readme = "README.md"
|
||||
requires-python = ">=3.13"
|
||||
dependencies = [
|
||||
"lxml>=6.0.4",
|
||||
]
|
||||
Reference in New Issue
Block a user