Files
CourseWork_IRFM/main.py
T
Dmitry 502f48a279 Add web crawler for legal documents
- Implemented a new crawler module to extract text and links from legal documents.
- Added a loader module to handle HTML downloads with error handling and retries.
- Created a main script to initiate the crawling process on a specified URL.
- Defined functions for normalizing URLs, extracting links, checking terminal pages, and extracting text content.
- Updated project metadata with dependencies and versioning in pyproject.toml and uv.lock.
2026-04-16 18:44:27 +03:00

9 lines
193 B
Python

from modules.crawler import crawl_document
def main():
url = "https://www.consultant.ru/document/cons_doc_LAW_10699/"
print(crawl_document(url))
if __name__ == "__main__":
main()