mirror of
https://github.com/alexhopeoconnor/bom-local-service.git
synced 2026-10-04 03:28:11 +10:00
Refactor scraping system to workflow-based architecture
Major architectural changes: - Replace monolithic ScrapingService with workflow-based system - Extract scraping logic into discrete, testable steps - Add configuration-driven selectors, JavaScript templates, and text patterns - Implement step registry and workflow factory for extensibility - Add PageState tracking and step dependencies - Make IWorkflow generic to support different response types New components: - Scraping steps organized by category (Navigation, Search, Map, Metadata, Capture) - BaseScrapingStep base class with common dependencies - ScrapingContext for shared state between steps - RadarScrapingWorkflow orchestrating step execution - SelectorService for configurable element selection - Configuration models (SelectorConfig, TextPatternsConfig, JavaScriptTemplatesConfig) Performance and monitoring: - Add step-level timing metrics with historical tracking - Log step durations and compare to averages - Warn on slow steps/workflows (50%+ slower than average) - Integrate metrics with CacheService Configuration improvements: - Externalize all selectors to appsettings.json - Externalize JavaScript templates to appsettings.json - Externalize text patterns to appsettings.json - Support Docker configuration via mounted appsettings.json - Remove hardcoded waits, use configurable timeouts Bug fixes: - Fix DebugService collection modification exception - Fix ResetToFirstFrameStep click interception with JavaScript fallback - Remove minutes ago fallback, use timestamp parsing only - Update timestamp parsing for new BOM website format Documentation: - Update README with new architecture details - Add configuration section with Docker guidance - Add development guidelines for extending scraping system - Add performance monitoring documentation
This commit is contained in:
@@ -0,0 +1,28 @@
|
||||
namespace BomLocalService.Services.Scraping;
|
||||
|
||||
/// <summary>
|
||||
/// Interface for a single scraping step
|
||||
/// </summary>
|
||||
public interface IScrapingStep
|
||||
{
|
||||
/// <summary>
|
||||
/// Unique name of the step
|
||||
/// </summary>
|
||||
string Name { get; }
|
||||
|
||||
/// <summary>
|
||||
/// Names of steps that must complete before this step can execute
|
||||
/// </summary>
|
||||
string[] Prerequisites { get; }
|
||||
|
||||
/// <summary>
|
||||
/// Checks if the step can execute in the current page state
|
||||
/// </summary>
|
||||
bool CanExecute(ScrapingContext context);
|
||||
|
||||
/// <summary>
|
||||
/// Executes the step
|
||||
/// </summary>
|
||||
Task<ScrapingStepResult> ExecuteAsync(ScrapingContext context, CancellationToken cancellationToken);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user