Mercurial > hg > egg-tcls
annotate fetch_feeds.tcl @ 422:880a07485275
Add utl_ctime() to utillib and use it elsewhere.
author | Matti Hamalainen <ccr@tnsp.org> |
---|---|
date | Sun, 08 Jan 2017 03:52:15 +0200 |
parents | 51c08336d7b1 |
children | 44c9128097cd |
rev | line source |
---|---|
0 | 1 #!/usr/bin/tclsh |
1 | 2 # |
3 # NOTICE! Change above path to correct tclsh binary path! | |
4 # | |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
5 ############################################################################## |
0 | 6 # |
323 | 7 # FeedCheck fetcher v1.0 by Matti 'ccr' Hamalainen <ccr@tnsp.org> |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
8 # (C) Copyright 2008-2015 Tecnic Software productions (TNSP) |
0 | 9 # |
10 # This script is freely distributable under GNU GPL (version 2) license. | |
11 # | |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
12 ############################################################################## |
0 | 13 |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
14 ### The configuration should be in config.feeds in same directory |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
15 ### as this script. Or change the line below to point where ever |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
16 ### you wish. See "config.feeds.example" for an example config file. |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
17 source [file dirname [info script]]/config.feeds |
0 | 18 |
422
880a07485275
Add utl_ctime() to utillib and use it elsewhere.
Matti Hamalainen <ccr@tnsp.org>
parents:
350
diff
changeset
|
19 ### Required utillib.tcl |
880a07485275
Add utl_ctime() to utillib and use it elsewhere.
Matti Hamalainen <ccr@tnsp.org>
parents:
350
diff
changeset
|
20 source [file dirname [info script]]/utillib.tcl |
880a07485275
Add utl_ctime() to utillib and use it elsewhere.
Matti Hamalainen <ccr@tnsp.org>
parents:
350
diff
changeset
|
21 |
0 | 22 |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
23 ############################################################################## |
139
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
24 |
0 | 25 package require http |
271 | 26 |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
27 if {[info exists http_user_agent] && $http_user_agent != ""} { |
296 | 28 ::http::config -urlencoding utf8 -useragent $http_user_agent |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
29 } else { |
322 | 30 ::http::config -urlencoding utf8 -useragent "Mozilla/5.0 (X11; Linux x86_64; rv:38.0) Gecko/20100101 Firefox/38.0" |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
31 } |
271 | 32 |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
33 if {[info exists http_use_proxy] && $http_use_proxy != 0} { |
63 | 34 ::http::config -proxyhost $http_proxy_host -proxyport $http_proxy_port |
0 | 35 } |
36 | |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
37 if {[info exists http_tls_support] && $http_tls_support != 0} { |
265
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
38 package require tls |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
39 ::http::register https 443 [list ::tls::socket -request 1 -require 1 -tls1 1 -cadir $http_tls_cadir] |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
40 } |
908edc54005a
feeds: Move configuration to separate file.
Matti Hamalainen <ccr@tnsp.org>
parents:
159
diff
changeset
|
41 |
0 | 42 |
268
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
43 ############################################################################## |
96310b1c88fa
feeds: Improve config resiliency.
Matti Hamalainen <ccr@tnsp.org>
parents:
265
diff
changeset
|
44 |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
45 proc fetch_dorequest { urlStr urlStatus urlSCode urlCode urlData urlMeta } { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
46 upvar 1 $urlStatus ustatus |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
47 upvar 1 $urlSCode uscode |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
48 upvar 1 $urlCode ucode |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
49 upvar 1 $urlData udata |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
50 upvar 1 $urlMeta umeta |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
51 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
52 if {[catch {set utoken [::http::geturl $urlStr -timeout 6000 -binary 1 -headers {Accept-Encoding identity}]} uerrmsg]} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
53 puts "HTTP request failed: $uerrmsg" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
54 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
55 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
56 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
57 set ustatus [::http::status $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
58 if {$ustatus == "timeout"} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
59 puts "HTTP request timed out ($urlStr)" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
60 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
61 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
62 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
63 if {$ustatus != "ok"} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
64 puts "Error in HTTP transaction: [::http::error $utoken] ($urlStr)" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
65 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
66 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
67 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
68 set ustatus [::http::status $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
69 set uscode [::http::code $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
70 set ucode [::http::ncode $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
71 set udata [::http::data $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
72 array set umeta [::http::meta $utoken] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
73 ::http::cleanup $utoken |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
74 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
75 return 1 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
76 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
77 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
78 |
139
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
79 proc add_entry {uname uprefix uurl utitle} { |
142 | 80 global currclock feeds_db nitems |
292
9f90d6918626
feeds: Also use the html entity conversion from utillib here.
Matti Hamalainen <ccr@tnsp.org>
parents:
271
diff
changeset
|
81 set utmp [utl_convert_html_ent $uurl] |
147 | 82 if {[string match "http://*" $utmp] || [string match "https://*" $utmp]} { |
83 set utest "$utmp" | |
84 } else { | |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
85 if {[string range $uprefix end end] != "/" && [string range $utmp 0 0] != "/"} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
86 set utest "$uprefix/$utmp" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
87 } else { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
88 set utest "$uprefix$utmp" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
89 } |
147 | 90 } |
139
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
91 |
296 | 92 set usql "SELECT title FROM feeds WHERE url='[utl_escape $utest]' AND feed='[utl_escape $uname]'" |
140
b0648e05c855
Change some variable names, etc.
Matti Hamalainen <ccr@tnsp.org>
parents:
139
diff
changeset
|
93 if {![feeds_db exists $usql]} { |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
94 # puts "NEW: $utest : $utitle" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
95 set usql "INSERT INTO feeds (feed,utime,url,title) VALUES ('[utl_escape $uname]', $currclock, '[utl_escape $utest]', '[utl_escape [utl_convert_html_ent $utitle]]')" |
142 | 96 incr nitems |
140
b0648e05c855
Change some variable names, etc.
Matti Hamalainen <ccr@tnsp.org>
parents:
139
diff
changeset
|
97 if {[catch {feeds_db eval $usql} uerrmsg]} { |
139
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
98 puts "\nError: $uerrmsg on:\n$usql" |
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
99 exit 15 |
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
100 } |
63 | 101 } |
0 | 102 } |
103 | |
104 | |
105 proc add_rss_feed {datauri dataname dataprefix} { | |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
106 if {[catch {set utoken [::http::geturl $datauri -binary 1 -timeout 6000 -headers {Accept-Encoding identity}]} uerrmsg]} { |
63 | 107 puts "Error getting $datauri: $uerrmsg" |
108 return 1 | |
109 } | |
110 set upage [::http::data $utoken] | |
111 ::http::cleanup $utoken | |
112 | |
113 set umatches [regexp -all -nocase -inline -- "<item>.\*\?<title><..CDATA.(.\*\?)\\\]\\\]></title>.\*\?<link>(http.\*\?)</link>.\*\?</item>" $upage] | |
114 set nmatches [llength $umatches] | |
115 for {set n 0} {$n < $nmatches} {incr n 3} { | |
116 add_entry $dataname $dataprefix [lindex $umatches [expr $n+2]] [lindex $umatches [expr $n+1]] | |
117 } | |
118 | |
119 if {$nmatches == 0} { | |
120 set umatches [regexp -all -nocase -inline -- "<item>.\*\?<title>(.\*\?)</title>.\*\?<link>(http.\*\?)</link>.\*\?</item>" $upage] | |
121 set nmatches [llength $umatches] | |
122 for {set n 0} {$n < $nmatches} {incr n 3} { | |
123 add_entry $dataname $dataprefix [lindex $umatches [expr $n+2]] [lindex $umatches [expr $n+1]] | |
124 } | |
125 } | |
0 | 126 |
63 | 127 if {$nmatches == 0} { |
128 set umatches [regexp -all -nocase -inline -- "<item \[^>\]*>.\*\?<title>(.\*\?)</title>.\*\?<link>(http.\*\?)</link>.\*\?</item>" $upage] | |
129 set nmatches [llength $umatches] | |
130 for {set n 0} {$n < $nmatches} {incr n 3} { | |
131 add_entry $dataname $dataprefix [lindex $umatches [expr $n+2]] [lindex $umatches [expr $n+1]] | |
132 } | |
133 } | |
143 | 134 |
63 | 135 return 0 |
0 | 136 } |
137 | |
138 | |
139 ############################################################################## | |
69
df3230f8aa46
Translate some comments to english and cosmetic fixes.
Matti Hamalainen <ccr@tnsp.org>
parents:
63
diff
changeset
|
140 ### Fetch and parse Halla-aho's blog page data |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
141 proc fetch_halla_aho { } { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
142 set datauri "http://www.halla-aho.com/scripta/"; |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
143 set dataname "Mestari" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
144 if {![fetch_dorequest $datauri ustatus uscode ucode upage umeta]} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
145 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
146 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
147 |
63 | 148 set umatches [regexp -all -nocase -inline -- "<a href=\"(\[^\"\]+\.html)\"><b>(\[^<\]+)</b>" $upage] |
149 set nmatches [llength $umatches] | |
150 for {set n 0} {$n < $nmatches} {incr n 3} { | |
151 add_entry $dataname $datauri [lindex $umatches [expr $n+1]] [lindex $umatches [expr $n+2]] | |
152 } | |
0 | 153 |
63 | 154 set umatches [regexp -all -nocase -inline -- "<a href=\"(\[^\"\]+\.html)\">(\[^<\]\[^b\]\[^<\]+)</a>" $upage] |
155 set nmatches [llength $umatches] | |
156 for {set n 0} {$n < $nmatches} {incr n 3} { | |
157 add_entry $dataname $datauri [lindex $umatches [expr $n+1]] [lindex $umatches [expr $n+2]] | |
158 } | |
0 | 159 } |
160 | |
161 | |
162 ### The Adventurers | |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
163 proc fetch_adventurers { } { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
164 set datauri "http://www.peldor.com/chapters/index_sidebar.html"; |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
165 set dataname "The Adventurers" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
166 if {![fetch_dorequest $datauri ustatus uscode ucode upage umeta]} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
167 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
168 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
169 |
63 | 170 set umatches [regexp -all -nocase -inline -- "<a href=\"(\[^\"\]+)\">(\[^<\]+)</a>" $upage] |
171 set nmatches [llength $umatches] | |
172 for {set n 0} {$n < $nmatches} {incr n 3} { | |
173 add_entry $dataname "http://www.peldor.com/" [lindex $umatches [expr $n+1]] [lindex $umatches [expr $n+2]] | |
174 } | |
0 | 175 } |
176 | |
177 | |
178 ### Order of the Stick | |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
179 proc fetch_oots { } { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
180 set datauri "http://www.giantitp.com/comics/oots.html"; |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
181 set dataname "OOTS" |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
182 if {![fetch_dorequest $datauri ustatus uscode ucode upage umeta]} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
183 return 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
184 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
185 |
63 | 186 set umatches [regexp -all -nocase -inline -- "<a href=\"(/comics/oots\[0-9\]+\.html)\">(\[^<\]+)</a>" $upage] |
187 set nmatches [llength $umatches] | |
188 for {set n 0} {$n < $nmatches} {incr n 3} { | |
189 add_entry $dataname "http://www.giantitp.com" [lindex $umatches [expr $n+1]] [lindex $umatches [expr $n+2]] | |
190 } | |
0 | 191 } |
192 | |
193 | |
350
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
194 ### Poliisi tiedotteet |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
195 proc fetch_poliisi { datauri dataname dataprefix } { |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
196 if {![fetch_dorequest $datauri ustatus uscode ucode upage umeta]} { |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
197 return 0 |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
198 } |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
199 |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
200 set umatches [regexp -all -nocase -inline -- "<div class=\"channelitem\"><div class=\"date\">(.*?)</div><a class=\"article\" href=\"(\[^\"\]+)\">(\[^<\]+)</a>" $upage] |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
201 set nmatches [llength $umatches] |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
202 for {set n 0} {$n < $nmatches} {incr n 4} { |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
203 set stmp [string trim [lindex $umatches [expr $n+3]]] |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
204 add_entry $dataname $dataprefix [lindex $umatches [expr $n+2]] "[lindex $umatches [expr $n+1]]: $stmp" |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
205 } |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
206 } |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
207 |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
208 |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
209 |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
210 |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
211 ### Open database, etc |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
212 set nitems 0 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
213 set currclock [clock seconds] |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
214 global feeds_db |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
215 if {[catch {sqlite3 feeds_db $feeds_dbfile} uerrmsg]} { |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
216 puts "Could not open SQLite3 database '$feeds_dbfile': $uerrmsg." |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
217 exit 2 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
218 } |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
219 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
220 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
221 ### Fetch the feeds |
350
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
222 fetch_poliisi "http://www.poliisi.fi/oulu/tiedotteet/1/0?all1/0" "Poliisi/Oulu" "http://www.poliisi.fi" |
51c08336d7b1
feeds: Add support for Poliisi.fi information reports.
Matti Hamalainen <ccr@tnsp.org>
parents:
323
diff
changeset
|
223 |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
224 fetch_halla_aho |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
225 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
226 fetch_adventurers |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
227 |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
228 fetch_oots |
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
229 |
143 | 230 #add_rss_feed "http://www.kaleva.fi/rss/145.xml" "Kaleva/Tiede" "" |
0 | 231 |
232 add_rss_feed "http://www.effi.org/xml/uutiset.rss" "EFFI" "" | |
233 | |
143 | 234 add_rss_feed "http://static.mtv3.fi/rss/uutiset_rikos.rss" "MTV3/Rikos" "" |
0 | 235 |
321
d8b957796121
feeds: Refactor the feeds fetching.
Matti Hamalainen <ccr@tnsp.org>
parents:
296
diff
changeset
|
236 #add_rss_feed "http://www.blastwave-comic.com/rss/blastwave.xml" "Blastwave" "" |
0 | 237 |
238 #add_rss_feed "http://lehti.samizdat.info/feed/" "Lehti" "" | |
239 | |
139
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
240 |
3305e142eecc
Change feed fetcher to use SQLite3 backend.
Matti Hamalainen <ccr@tnsp.org>
parents:
114
diff
changeset
|
241 ### Close database |
140
b0648e05c855
Change some variable names, etc.
Matti Hamalainen <ccr@tnsp.org>
parents:
139
diff
changeset
|
242 feeds_db close |
142 | 243 |
244 puts "$nitems new items." |