forked from darshanime/notes
-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathpython.org
More file actions
executable file
·3186 lines (2224 loc) · 104 KB
/
Copy pathpython.org
File metadata and controls
executable file
·3186 lines (2224 loc) · 104 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
960
961
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
995
996
997
998
999
1000
* python
1. Tuples are accessible like lists.
a = (5, 4)
print a[0] --> 5
create a tuple with one entry
b = a,
print type(b) --> tuple
Tuples are immutable, like dict's keys.
tuples are hashable - they can be passed to hash() which is a function from the __builtin__ module. This is because they are non-mutable.
Use it like this : print __builtins__.hash(tup)
Although it is not necessary, it is conventional to enclose tuples in parentheses
so, this is valid as well:
a = 1, 2, 3, 4, 5
like strings, tuples are immutable. Once Python has created a tuple in memory, it cannot be changed.
A tuple lets us “chunk” together related information and use it as a single thing.
a = 1, 2, 3, 4
a is a tuple
b, c, d, e = a
b, c, d and e are ints
b = ("Bob", 19, "CS") # tuple packing
(name, age, studies) = b # tuple unpacking
The tuple can be heterogenous as well - for eg, it can have string, int, list.
a = 1, 2, 3, range(4), 6
if it has a list, it cant be hashed because one element - the list, is mutable
Look at this to add items to tuples from lists etc
b = [1, 2, 3]
temp = ()
lis = list(temp)
for x in b:
lis.append(int(x))
tup = tuple(lis)
print __builtins__.hash(tup)
2. take care of types of items. a=[]
a.append("1")
a.index(1) --> ValueError, not found.
Because "1" is stored, not 1.
WRONG^
because we are not storing anything at index 1, we have "1" at index 0
3. print("%.2f" % a)
to print floats to 2 decimal places.
4. SETS
Concept
If the inputs are given in one line separated by a space character, use split() to get the splitted values in form of a list. EG:
a = raw_input()
5 4 3 2
lis = a.split()
print (lis)
['5', '4', '3', '2']
If the values in a list are all of integer type, use the map() to convert all the strings to integers.
map(function_object_to_use_on_each_item, iterable)
ALSO: lambda <variable it takes> : <variable it returns>
so, lambda x: (x, x**2, x**3)
means, it takes in x and returns a tuple of values
we can make lambda take in 2 args as well:
map(lambda x, y: (x**2, y**2), range(10), range(10))
here, we will get a tuple of 10 elements, each will have int squared from 0 to 9.
map(str, range(10)) - works
map(str, range(10), range(10)) - str doesnt accept two args, only 1, 2 given
lambda x: (x, x**2, x**3)
<function <lambda> at 0x7f1688c35aa0>
Store this function object and call it with new values
x = lambda x: (x, x**2, x**3)
x(3)
(3, 9, 27)
We could use x as the variable name to store lambda function because of namespaces. The function namespace is different from the outside one.
newlis = list(map(int, lis))
print (newlis)
[5, 4, 3, 2]
Sets are unordered bag of unique values. A single set contains values of any immutable data type.
A set can't store immutable data type.
so, s = {1, 2, 3, range(10)} is not allowed
sets can be updated, so they arent hashable
CREATING SET
myset = {1, 2} # Directly assigning values to a set
myset = set() # Initializing a set
myset = set(['a', 'b']) # Creating a set from a list
myset
{'a', 'b'}
MODIFYING SET - add() and update()
myset.add('c')
myset
{'a', 'c', 'b'}
myset.add('a') # As 'a' already exists in the set, nothing happens
myset.add((5, 4))
myset
{'a', 'c', 'b', (5, 4)}
myset.update([1, 2, 3, 4]) # update() only works for iterable objects
myset
{'a', 1, 'c', 'b', 4, 2, (5, 4), 3}
myset.update({1, 7, 8})
myset
{'a', 1, 'c', 'b', 4, 7, 8, 2, (5, 4), 3}
myset.update({1, 6}, [5, 13])
myset
{'a', 1, 'c', 'b', 4, 5, 6, 7, 8, 2, (5, 4), 13, 3}
REMOVING ITEMS - discard() and remove()
Both discard() and remove() take a single value as an argument and removes that value from the set. If that value is not present in the set, discard() does nothing but remove() raises a KeyError exception
myset.discard(10)
myset
{'a', 1, 'c', 'b', 4, 5, 7, 8, 2, 12, (5, 4), 13, 11, 3}
myset.remove(13)
myset
{'a', 1, 'c', 'b', 4, 5, 7, 8, 2, 12, (5, 4), 11, 3}
COMMON SET OPERATIONS - union(), intersection() and difference()
a = {2, 4, 5, 9}
b = {2, 4, 11, 12}
a.union(b) # Values which exist in a or b
{2, 4, 5, 9, 11, 12}
a.intersection(b) # Values which exist in a and b
{2, 4}
a.difference(b) # Values which exist in a but not in b
{9, 5}
union() and intersection() are symmetric methods i.e. to say,
a.union(b) == b.union(a)
True
a.intersection(b) == b.intersection(a)
True
a.difference(b) == b.difference(a)
False
5. MAP function
a = map(X, Y)
X is a function. Can be str, int, lambda x : x**2
Y is a the input on which to apply the function. can be iterable.
a = range(10)
print map(lambda x : x**2, a)
6.raw_input() ALWAYS INPUTS A STR. CONVERT TO INT IF NEEDED.
join takes in an interable of STRINGS only and returns a single string.
join is a method of the String class, it can act on string objects only
so, "-".join(range(10))
doesnt work
but
"-".join(map(str, range(10)))
does
7. a = raw_input()
print a.split()
print "-".join(["hello", "i", "am", "dc"])
or
print a.replace(" ", "-")
8. STRING MANIPULATION :
>>> string = "abracadabra"
>>> l = list(string)
>>> l[5] = 'k'
>>> string = ''.join(l)
or
string = string[:5] + "k" + string[6:]
9. STRING MANIPULATION :
print str1+str2
str1.upper(), str1.lower(), str1.swapcase(), str1.capitalize() #only 1st letter of string will be CAPSed
print str1[1:5]
str1.find('llo') # find the index from which the first instance of substr llo begins.If not found, -1
str1.rfind('l') # find the index of 'l' but start from reverse - finds the last occurance of l
str1.replace('l', 'r') # replaces ALL occurances
str1.strip() #strips the whitespaces
str1.isalnum() # is alpha-numerical eg ab123
str1.isalpha() # is aplha eg abcD but not ab12
str1.isdigit() # is digit, eg 123, not 123a
str1.islower()
str1.isupper()
str1.rjust/ljust/center(int for width, #optional "-" - what to fill the remaining space with, default is whitespace)
print str1*25 #will print it 25 times.
10. ANY FUNCTION
Python has a function called any() that returns True if any one of the list elements evals to True.
takes in an iterable and returns a boolean
ex:
print(any([0, 1, 0, 0])) # will print True
print(any([0, 0, 0, 0])) # will print False
11. REDUCE FUNCTION :
>>> f = lambda a,b: a if (a > b) else b #IF ELSE IN LAMBDA
>>> reduce(f, [47,11,42,102,13]) # APPLIED TO FIRST 2 ELEMENTS, THEN THE RESULT+THE THIRD ELEMENT
eg : sum of the first 100 elements
print reduce(lambda x,y:x+y, range(1,101))
At first the first two elements of seq will be applied to func, i.e. func(s1,s2) The list on which reduce() works looks now like this: [ func(s1, s2), s3, ... , sn ]
In the next step func will be applied on the previous result and the third element of the list, i.e. func(func(s1, s2),s3)
The list looks like this now: [ func(func(s1, s2),s3), ... , sn ]
Continue like this until just one element is left and return this element as the result of reduce()
REDUCE RETURNS ONE VALUE IN THE END
12. BOOL()
print bool(1) #TRUE
print bool("a") # TRUE
print bool(0) #FALSE
print bool("0") #TRUE - because it is a string
13. TEXTWRAP :
>>> import textwrap
>>> string = "This is a very very very very very long string."
>>> print textwrap.wrap(string,8)
['This is', 'a very', 'very', 'very', 'very', 'very', 'long', 'string.']
Returns a list of strings of given size - it breaks down the very big string.
>>> import textwrap
>>> string = "This is a very very very very very long string."
>>> print textwrap.fill(string,8)
Prints a single string with each line not more than the specied width.
14. RANGE/XRANGE
print range(1,10,2)
[1, 3, 5, 7, 9]
print range(10, 1, -2)
[10, 8, 6, 4, 2]
15. NEW VARIANT OF DICT
from collections import defaultdict
d = defaultdict(list) #YOU HAVE TO PREDEFINE THE DATATYPE OF THE DICT'S VALUES FIELD
d['python'].append("awesome")
d['something-else'].append("not relevant")
d['python'].append("language")
for i in d.items():
print i
16. THIS IS THE CODE FOR THE NO IDEA CHALLENGE
from collections import defaultdict
d=defaultdict(list)
n_n, n_ab = map(int, raw_input().strip().split(' '))
n = map(lambda x : d[x].append(1), raw_input().strip().split(' '))
a = map(str, raw_input().strip().split(' '))
b = map(str, raw_input().strip().split(' '))
h=0
for i in xrange(n_ab):
print a[i], d[a[i]]
if d[a[i]] != []:
h+=sum(d[a[i]])
if d[b[i]] != []:
h-=sum(d[b[i]])
print h
When you wish to count the occurances of an item in a big array and manipulate it later, use dict. the key is that item and the value is a list appended by 1 (so, you can sum the values to find #of occurrences) or the index etc. - for eg if it is given in lines.
Take a look at :
# Enter your code here. Read input from STDIN. Print output to STDOUT
from collections import defaultdict
d = defaultdict(list)
n,m=map(int,raw_input().strip().split(' '))
for i in xrange(1,n+1):
s=raw_input().strip()
d[s].append(i)
for i in xrange(m):
s=raw_input().strip()
if d[s]!=[]:
print " ".join(map(str,d[s]))
else:
print "-1"
17. PRINT LIST ON THE SAME LINE
a = range(10)
print a - [0, 2, ..., 9]
but
for i in a:
print a
will give :
0
1
2
3
..
9
For : 0, 1, 2, .., 9 do print a, or print (a, end=" ") #PYTHON-3
17. TIP
Sometime when timing out even with the correct code, sit back and relaize how you solved the problem.
1. Storing millions of values is not a problem
2. Use xrange and never range
3. The time consuming task are the LOOPS. If you have to traverse the many times, it can be a problem.
Think about the various scenarios and try to figure out a means to simplify the problem. There is a trick, you just need to crack it.
18. There is deque() to replace list. It can act as a stack, queue etc. Very fast.
19. Set is unordered collection, cannot have duplicate entries.
a = set()
set([1, 1,2, 3])
-- will store only one one
a = dict
print set(a) ##--will print the unique keys present in a
SETS ARE GENERALLY USED FOR MEMBERSHIP TESTING AND DUPLICATE ENTRIES ELIMINATING
a=set('HackerRank')
a.add('H') ##-- returns none. so print a.add('H') will print: `None`
SETS : DIFFERENCE BETWEEN REMOVE AND DISCARD
.remove(x)
This operation removes element x from set.
If element x is not in the set, it raises a KeyError.
.remove(x) operation returns None
.discard(x)
This operation also removes element x from set.
But if element x is not in the set, it does not raises a KeyError.
.discard(x) operation returns None.
.pop()
This operation removes and return an arbitrary element from set.
If there are no elements to remove, it raises a KeyError.
.union()
.union() operator returns the union of set and the set of elements in an iterable.
Sometimes '|' operator is used in place of .union() operator but it operates only on the set of elements in set.
Set is immutable to .union() operation (or '|' operation).
>>> s = set("Hacker")
>>> print s.union("Rank" OR DICT OR LIST OR TUPLES OR ENUMERATE(LISTS) ETC)
>>> s | set("Rank") # ANOTHER WAY TO WRITE ABOUT IT
CHAINING COMMANDS IS POSSIBLE ONLY IF THE INSTANCE RETURNED IS COMPATIBLE
EXAMPLE : str1.strip().split(" ") - is possible because strip will return str, split will return list.
NOW, IN SETS :
req = set()
req.update(set2).update(set23)
is not allowed because the first update returns a NONE, and AttributeError: 'NoneType' object has no attribute 'update'
.intersection()
.intersection() operator returns the intersection of set and the set of elements in an iterable.
Sometimes '&' operator is used in place of .intersection() operator but it operates only on the set of elements in set.
Set is immutable to .intersection() operation (or '&' operation).
.difference()
.difference() returns a set with all elements from set that are not in an iterable.
Sometimes '-' operator is used in place of .difference() operator but it operates only on the set of elements in set.
Set is immutable to .difference() operation (or '-' operation).
20. THERE ARE TWO TYPES OF METHODS USED TO ALTER THE OBJECT.
str1.replace(" ", "-")
and list.sort()
NOW THE FORMER RETURNS A STR AND YOU CAN PRINT IT ETC. BUT IT DOESNT CHANGE STR1. STR1 STILL HAS SPACES AND NOT DASHES.
WHEREAS THE LATTER RETURNS NOTHING AND MODIFIES THE LIST IN-PLACE.
NOWHERE IS IT POSSIBLE THAT THE SAME FUNCTION CALL MUTATES THE OBJECT, AND RETURNS THE MUTATED OBJECT.
so, you can either copy the object, change it and return it like by replace
or you can modify it in place and return nothing
21. FOR DEALING WITH COMPLEX NUMBERS, USE CMATH MODULE
from cmath import phase
print phase(complex(-1, 0)) --> 3.141...
22. CARTESIAN PRODUCT IS A MATHEMATICAL OPERATION ACC TO WHICH EACH ELEMENT FROM A LIST IS OPERATED ALONG WITH EACH ELEMENT FROM THE OTHER SET.
AxB = [(a,b) for each a belonging to A and each b belonging to B]
PYTHON :
PRINT [(a, b) FOR a in A for b in B]
SAME THING IS DONE USINT ITERTOOLS
FROM ITERTOOLS IMPORT PRODUCT
PRODUCT(A, B)
23. itertools.permutations(iterable[, r])
Returns successive r length permutations of elements in an iterable.
If r is not specified or is None, then r defaults to the length of the iterable and all possible full-length permutations are generated.
Permutations are emitted in lexicographic sort order. So, if the input iterable is sorted, the permutation tuples will be produced in sorted order.
<itertools.product object at 0x7f00e09d4f00>
THIS WILL BE PRINTED WHEN YOU PRINT DIRECTLY : PRINT PRODUCT(A, B)
TO ACTUALLY ITERATE THEM, ENCLOSE THEM IN A LIST EG : LIST(PRODUCT(A,B))
24. itertools.combinations(iterable, r)
Return r length subsequences of elements from the input iterable.
Combinations are emitted in lexicographic sort order. So, if the input iterable is sorted, the combination tuples will be produced in sorted order.
>>> from itertools import combinations
>>>
>>> print list(combinations('12345',2))
[('1', '2'), ('1', '3'), ('1', '4'), ('1', '5'), ('2', '3'), ('2', '4'), ('2', '5'), ('3', '4'), ('3', '5'), ('4', '5')]
>>>
>>> A = [1,1,3,3,3]
>>> print list(combinations(A,4))
[(1, 1, 3, 3), (1, 1, 3, 3), (1, 1, 3, 3), (1, 3, 3, 3), (1, 3, 3, 3)]
25. THERE IS A CERTAIN PROCEDURE OF THINKING ABOUT HOW TO SOLVE THE PROBLEM :
i) THINK ABOUT THE DATATYPE TO USE TO STORE THE INPUT - LIST/DICT/TUPLE/SET ETC.
ii) ACCEPT THE DATA AND STORE THEM PROPERLY.
iii) APPLY THE LOGIC AND GET THE REQUIRED RESULT
iv) MANIPULATE THE DATATYPE HOLDING THE RESULT AND DISPLAY IT IN THE REQUIRED WAY EG USE "".JOIN(LIST1) ETC.
26. collections.Counter()
A counter is container, where elements are stored as dictionary keys and their counts are stored as dictionary values.
Sample Code
>>> from collections import Counter
>>>
>>> myList = [1,1,2,3,4,5,3,2,3,4,2,1,2,3]
>>> print Counter(myList)
Counter({2: 4, 3: 4, 1: 3, 4: 2, 5: 1})
>>>
>>> print Counter(myList).items()
[(1, 3), (2, 4), (3, 4), (4, 2), (5, 1)]
>>>
>>> print Counter(myList).keys()
[1, 2, 3, 4, 5]
>>>
>>> print Counter(myList).values()
[3, 4, 4, 2, 1]
27.
import calendar
>>>
>>> print calendar.TextCalendar(firstweekday=6).formatyear(2015)
2015
January February March
Su Mo Tu We Th Fr Sa Su Mo Tu We Th Fr Sa Su Mo Tu We Th Fr Sa
1 2 3 1 2 3 4 5 6 7 1 2 3 4 5 6 7
4 5 6 7 8 9 10 8 9 10 11 12 13 14 8 9 10 11 12 13 14
11 12 13 14 15 16 17 15 16 17 18 19 20 21 15 16 17 18 19 20 21
18 19 20 21 22 23 24 22 23 24 25 26 27 28 22 23 24 25 26 27 28
25 26 27 28 29 30 31 29 30 31
28
>>> import string
>>> string.ascii_lowercase
'abcdefghijklmnopqrstuvwxyz'
list(string.ascii_lowercase)
29
SORTING LISTS BY MULTIPLE KEYS
a = [('a', 3), ('a', 2), ('b', 4), ('c', 5)]
print sorted(a, key=lambda d : (d[0], -d[1]))
sorted(<iterable>, key=<function that takes in each element of the iterable and returns tuple - the first entry is tried to sort, in case of ties, second entry is tried>)
sorted in increasing order wrt to the keys
30
zip([iterable, ...])
This function returns a **list of tuples**, where the i-th tuple contains the i-th element from each of the argument sequences or iterables.
If argument sequences are of unequal lengths, then returned list is truncated in length to the length of the shortest argument sequence.
31
A = [1,2,3]
B = [6,5,4]
C = [7,8,9]
X = A + B + C
print X
[1, 2, 3, 6, 5, 4, 7, 8, 9]
X = [A]+[B]+[C]
print X
[[1, 2, 3], [6, 5, 4], [7, 8, 9]]
32
ZeroDivisionError
Raised when the second argument of a division or modulo operation is zero.
ValueError
Raised when a built-in operation or function receives an argument that has the right type but an inappropriate value.
try and except statements can be used to handle selected exceptions. A try statement may have more than one except clause, to specify handlers for different exceptions.
try:
print 1/0
except ZeroDivisionError as e:
print "Error Code:",e
#Output
Error Code: integer division or modulo by zero
33
Concept
The map() function applies a function to every member of an iterable and returns the result. It takes two parameters, first the function which is to be applied and second the iterables like a list.
Let's say you are given a list of names and you have to print a list which contains length of each name.
>> print (list(map(len, ['Tina', 'Raj', 'Tom'])))
[4, 3, 3]
Lambda is a single expression anonymous function often used as an inline function. In simple words, it is a function which has only one line in its body. It proves very handy in functional and GUI programming.
>> sum = lambda a, b, c: a + b + c
>> sum(1, 2, 3)
6
Note:
Lambda functions cannot use the return statement and can only have a single expression. Unlike def, which creates a function and assigns it a name, lambda creates a function and returns the function itself. Lambda can be used inside list and dictionary.
34
**The re.sub() (sub stands for substitution) evaluates a pattern and for each valid match, it calls a method (or lambda).**
SO, RE.SUB() TAKES 3 ARGUEMENTS. THE REGEX, THE FUNCTION/LAMDBA TO APPLY TO THE MATCHES AND THE STRING
EXAMPLE 1 :
print map(lambda x:x, "1 2 3 4 5")
['1', ' ', '2', ' ', '3', ' ', '4', ' ', '5']
^^HERE, THE STRING IS `LIST`-ED AND EVERY ELEMENT IS GIVEN TO LAMBDA WHICH JUST RETURNS IT.
NOW,
EXAMPLE 2 :
print re.sub(r"\d+", lambda x:x, "1 2 3 4 5")
The method is called for all matches and can be used to modify strings in different ways.
The re.sub() method returns the modified string as an output.
import re
#Squaring numbers
def square(match):
number = int(match.group(0))
return str(number**2)
print re.sub(r"\d+", square, "1 2 3 4 5 6 7 8 9")
35 VALID EMAIL ID : x IS THE STR VAR CONTAINING THE EMAIL ID
re.findall('([\w-]+)@([a-z0-9]+)\.([\w]+)', x)
36. LISTS GYAN
If both slice indices are left out, all items of the list are included. But this is not the same as the original a_list variable. It is a new list that happens to have all the same items. a_list[:] is shorthand for making a complete copy of a list.
a = range(3)
id(a)==id(a[:])
False
Slicing works if one or both of the slice indices is negative. If it helps, you can think of it this way: reading the list from left to right, the first slice index specifies the first item you want, and the second slice index specifies the first item you don’t want. The return value is everything in between.
WHEN PRINTING, IF THE START INDEX IS TO THE RIGHT OF THE END INDEX, NOTHING IS PRINTED.
EG :
a = range(100)
print a[2:4]
[2, 3]
print a[5:2]
[]
print a[-4:5]
[]
print a[-5:-1]
[95, 96, 97, 98]
+ OPERATOR ADDS A LIST TO THE EXISTING LIST
The append() method adds a single item to the end of the list.
The insert() method inserts a single item into a list. The first argument is the index of the first item in the list that will get bumped out of position. EG: A_LIST.INSERT(0, 'HI')
APPEND VS EXTEND
The extend() method takes a single argument, which is always a list, and adds each of the items of that list to a_list.
>>> a_list = ['a', 'b', 'c']
>>> a_list.extend(['d', 'e', 'f']) ①
>>> a_list
['a', 'b', 'c', 'd', 'e', 'f']
>>> a_list.append(['g', 'h', 'i']) ③
>>> a_list
['a', 'b', 'c', 'd', 'e', 'f', ['g', 'h', 'i']]
37
SEARCHING IN LISTS
>>> a_list = ['a', 'b', 'new', 'mpilgrim', 'new']
>>> a_list.count('new') ①
2
>>> 'new' in a_list ②
True
>>> 'c' in a_list
False
>>> a_list.index('mpilgrim') ③
3
>>> a_list.index('new') ④
2
>>> a_list.index('c') ⑤
Traceback (innermost last):
File "<interactive input>", line 1, in ?
ValueError: list.index(x): x not in list
COUNT() - RETURNS THE COUNT OF THE ITME IN LIST
`IN` - TELLS YOU IF ITEM IN THE LIST OR NOT
`INDEX` - TELLS YOU WHERE IN THE LIST IS THE ITEM. IF NOT THERE, VALUeERROR
REMOVE ITEMS FROM THE LIST: DEL A_LIST[1]
OR A_LIST.REMOVE('HELLO') - REMOVES THE FIRST INSTANCE OF HELLO ONLY
A_LIST.POP() - REMOVES THE LAST ITEM AND RETURNS IT.
You can pop arbitrary items from a list. Just pass a positional index to the pop() method. It will remove that item, shift all the items after it to “fill the gap,” and return the value it removed.
IN BOOLEAN CONTEXT, EMPTY LIST IS FALSE. OTHERS ARE TRUE
TUPLES
A tuple is defined in the same way as a list, except that the whole set of elements is enclosed in parentheses instead of square brackets.
The elements of a tuple have a defined order, just like a list. Tuple indices are zero-based, just like a list, so the first element of a non-empty tuple is always a_tuple[0]
SLICING WORKS, IT RETUENS A NEW TUPLE.
The major difference between tuples and lists is that tuples can not be changed. In technical terms, tuples are immutable. In practical terms, they have no methods that would allow you to change them. Lists have methods like append(), extend(), insert(), remove(), and pop(). Tuples have none of these methods.
TUPLES HAVE A_TUPLE.INDEX('HELLO') AND 'HELLO' IN A_TUPLE
So what are tuples good for?
Tuples are faster than lists. If you’re defining a constant set of values and all you’re ever going to do with it is iterate through it, use a tuple instead of a list.
It makes your code safer if you “write-protect” data that doesn’t need to be changed. Using a tuple instead of a list is like having an implied assert statement that shows this data is constant, and that special thought (and a specific function) is required to override that.
Some tuples can be used as dictionary keys (specifically, tuples that contain immutable values like strings, numbers, and other tuples). Lists can never be used as dictionary keys, because lists are not immutable.
☞Tuples can be converted into lists, and vice-versa. The built-in tuple() function takes a list and returns a tuple with the same elements, and the list() function takes a tuple and returns a list. In effect, tuple() freezes a list, and list() thaws a tuple.
To create a tuple of one item, you need a comma after the value. Without the comma, Python just assumes you have an extra pair of parentheses, which is harmless, but it doesn’t create a tuple.
EG : A = (1, )
RANGE() RETURNS AN ITERATOR NOT A LIST/TUPLE
38. RETURN MULTIPLE ITEMS FROM A FUNCTION
You can also use multi-variable assignment to build functions that return multiple values, simply by returning a tuple of all the values. The caller can treat it as a single tuple, or it can assign the values to individual variables.
39 SETS
A set is an unordered “bag” of unique values. A single set can contain values of any immutable datatype. Once you have two sets, you can do standard set operations like union, intersection, and set difference.
SO:
Lists - mutable - can contain mutable datatypes - can't be hashed - ordered
Tuples - immutable - can contain mutable datatypes - can be hashed if they contain no mutable datatype -
sets - mutable - cannot contain mutable datatype - cannot be hashed - unordered
dicts - mutalbe - can contain mutable datatypes(not as keys but) - can't be hashed - unorder (Ordereddict is ordered)
CREATE A NEW SET :
A_SET = {1}
CREATE A EMPTY SET :
A_SET = SET()
sets can hold UNMUTABLE DATATYPES ONLY. SO NO LISTS IN SETS. TUPLES ALLOWED.
The update() method takes one argument, a set, and adds all its members to the original set. It’s as if you called the add() method with each member of the set.
② Duplicate values are ignored, since sets can not contain duplicates.
③ You can actually call the update() method with any number of arguments. When called with two sets, the update() method adds all the members of each set to the original set (dropping duplicates).
④ The update() method can take objects of a number of different datatypes, including lists. When called with a list, the update() method adds all the items of the list to the original set.
40. REMOVE DATA FROM SETS
1. REMOVE() - IF ELEMENT NOT PRESENT IN SET, RAISE ERROR
2. DISCARD() - IF ELEMENT NOT PRESENT, DO NOT RAISE ERROR
3. POP() - RETURNS A RANDOM VALUE - BCOZ SETS ARE UNORDERED
4. CLEAR() - REMOVES ALL VALUES FROM THE SET
41. COMMON SET OPERATIONS:
1. 'A' IN A_SET - RETURNS BOOLEAN - TRUE/FALSE
2. A_SET.UNION/INTERSECTION/DIFFERENCE/SYMMETRIC_DIFFERENCE(B_SET)
UNION - RETURNS A NEW SET HAVING ALL ELEMENTS OF BOTH A AND B
INTERSECTION - BOTH SETS
DIFFERENCE - IN A BUT NOT IN B : A-B - NOT A SYMMETRIC OPERATION
SYMMETRIC_DIFFERENCE - ONLY ONCE IN EITHER A OR B
41. EXTRA OPERATIONS ON SETS
A_SET.ISSUBSET(B_SET)
A_SET.ISSUPERSET(B_SET)
41. 'HELLO' IN A_DICT - WILL RETURN TRUE IF 'HELLO' IS A KEY OF THE DICT
41. NONE IS SPEACIAL. IT IS NOT 0, FALSE, EMPTY ETC
NONE IS NULL
NONE==NONE TRUE, ELSE ALWAYS FALSE
NONE EVALUATES TO FALSE AND not NONE TO TRUE
42. OS MODULE
OS.GETCWD()
OS.CHDIR() - CHANGES THE CURRENCT WORKING DIR
OS.PATH - CONTAINS FUNCTIONS FOR MANIPULATING FILENAMES AND DIR NAMES
OS.PATH.JOIN() - TAKES TWO OR MORE PARTIAL FILEPATHS AND MAKES THEM ONE VALID PATHNAME AUTOMATICALLY BASED ON YOUR OS.
OS.PATH.EXPANDUSER() - EXPANDS A PATHNAME THAT USES ~ TO REPRESENT THE CURRENT USER'S HOME DIR.
OS.PATH.SPLIT(PATHNAME) - SPLITS THE PATH AND FILENAME SEPERATELY
OS.PATH.SPLITTEXT(FILENAME) - SPLITS THE FILENAME AND IT'S EXTENSION
43. GLOB
SPECIALITY IS THAT IT ACCEPTS WILDCARDS
GLOB.GLOB('EXAMPLES/*.MP3')
44. METADATA ABOUT THE FILE :
LIKE SIZE, TIME OF CREATION ETC.
metadata = os.stat('hello.py')
metadata.st_mtime - MODIFICATION TIME
--> will print the time? - THE NUMBER OF SECS SINCE THE EPOCH - JAN1, 1970
metadata.st_size
- will be in bytes
import humansize - converts bytes to human readable form.
humansize.approximate_size(metadata.st_size)
3.1 KiB
^THE SAME BLOB OF NUMBER LIKE WE HAD FOR FACE DETECTION. USE : TIME.LOCALTIME(`THAT INT`) TO GET THE TIME, DATE ETC
45. GET ABS PATH OF A FILE
OS.PATH.REALPATH('HELLO.PY')
46. DICTIONARY COMPREHENSIONS :
JUST LIKE LIST COMPREHENSIONS, BUT CREATE A DICT AND NOT A LIST
a = [i**2 for i in range(10)]
a is a list
a = {i:i**2 for i in range(10)}
a is a dict above.
REPLACE KEYS AND VALUES IN DICT
a = {value:key for key, value in a_dict}
^wont work if the values are lists. because lists cannot be keys to any dict as they are immutable.
47. SET COMPREHENSIONS
A_SET = SET(RANGE(10))
B_SET = {X**2 FOR X IN A_SET}
48. EACH CHARACTER IS ENCODED DIFFERENTLY. FOR EXAMPLE, THE CHAR `A` IS STORED DIFFERENTLY IN MEMORY IN THE ASCII FORMAT, UTF-8 ETC. TO GET BACK THE A, YOU NEED THE KEY - THAT IS YOU NEED TO KNOW IN WHAT WAY TO INTEREPET THE DATA.
EXAMPLES OF ENCODINGS :
ASCII - STORES ENGLISH CHARACTERS AS NUMBERS RANGING FROM 0 TO 127
65 IS A, 97 IS a ETC.
PLAIN TEXT IS WHAT YOU WRITE ON PAPER. EG: 'hello'
THIS IS ENCODED TO BYTES IN A PARTICULAR WAY ACC TO THE CHARACTER ENCODING.
ENTER UNICODE
Unicode is a system designed to represent every character from every language. Unicode represents each letter, character, or ideograph as a 4-byte number. Each number represents a unique character used in at least one of the world’s languages. There is exactly 1 number per character, and exactly 1 character per number. Every number always means just one thing; there are no “modes” to keep track of. U+0041 is always 'A', even if your language doesn’t have an 'A' in it.
THAT IS CALLED UTF-32 (32 BITS = 4 BYTES)
THEN THERE IS UTF-16 (2 BYTES FOR EACH CHAR)
UTF-8 (VARIALBE LENGTH ENCODING SYSTEM) - FOR ASCII - JUST ONE BYTE USED
In Python 3, all strings are sequences of Unicode characters. There is no such thing as a Python string encoded in UTF-8, or a Python string encoded as CP-1252. “Is this string UTF-8?” is an invalid question. UTF-8 is a way of encoding characters as a sequence of bytes. If you want to take a string and turn it into a sequence of bytes in a particular character encoding, Python 3 can help you with that. If you want to take a sequence of bytes and turn it into a string, Python 3 can help you with that too. Bytes are not characters; bytes are bytes. Characters are an abstraction. A string is a sequence of those abstractions.
49
>>> username = 'mark'
>>> password = 'PapayaWhip' ①
>>> "{0}'s password is {1}".format(username, password) ②
"mark's password is PapayaWhip"
0 REFERS TO THE FIRST ARGUMENT PASSED TO FORMAT.
IF A USERNAME IS A LIST: 0[0] WOULD BE THE FIRST ELEMENT
THIS WORKS TOO :
>>> import humansize
>>> import sys
>>> '1MB = 1000{0.modules[humansize].SUFFIXES[1000][0]}'.format(sys) #NORMALLY, YOU WOULD PUT QUOTES AROUND humansize BECAUSE THAT KEY IS A STR. BUT HERE, IT IS NOT required
'1MB = 1000KB'
50. SYS MODULE
SYS MODULE STORES INFORMATION ABOUT THE CURRENTLY RUNNIG PYTHON INSTANCE
SYS.MODULES - LIST OF ALL THE MODULES IMPORTED INTO PYTHON
51 BYTES
Bytes are bytes; characters are an abstraction. An immutable sequence of Unicode characters is called a string. An immutable sequence of numbers-between-0-and-255 is called a bytes object.
To define a bytes object, use the b'' “byte literal” syntax. Each byte within the byte literal can be an ASCII character or an encoded hexadecimal number from \x00 to \xff (0–255).
② The type of a bytes object is bytes.
③ Just like lists and strings, you can get the length of a bytes object with the built-in len() function.
④ Just like lists and strings, you can use the + operator to concatenate bytes objects. The result is a new bytes object.
52. DEFAULT ENCODING
Python 3 assumes that your source code — i.e. each .py file — is encoded in UTF-8.
☞In Python 2, the default encoding for .py files was ASCII. In Python 3, the default encoding is UTF-8.
If you would like to use a different encoding within your Python code, you can put an encoding declaration on the first line of each file. This declaration defines a .py file to be windows-1252:
# -*- coding: windows-1252 -*-
Technically, the character encoding override can also be on the second line, if the first line is a UNIX-like hash-bang command.
#!/usr/bin/python3
# -*- coding: windows-1252 -*-
53. SIMPLE REPLACE BY STRINGS
STR_.REPLACE("HELLO", "HI")
IF YOU NEED POWERFUL REGEX AIDED REPLACEMENT
RE.SUB(REGEXpATTER, REPLR_STR, string)
54. REGEX EXAMPLES
>>> pattern = '^M?M?M?(CM|CD|D?C?C?C?)$' ①
>>> re.search(pattern, 'MCM') ②
<_sre.SRE_Match object at 01070390>
>>> re.search(pattern, 'MD') ③
<_sre.SRE_Match object at 01073A50>
>>> re.search(pattern, 'MMMCCC') ④
<_sre.SRE_Match object at 010748A8>
>>> re.search(pattern, 'MCMC') ⑤
>>> re.search(pattern, '') ⑥
<_sre.SRE_Match object at 01071D98>
'^M?M?M?$' - THIS SAYS THERE ARE 0-3 M'S THAT WOULD BE ACCEPTED. SO, M/MM/MMM WOULD GO IN
BETTER WAY TO EXPRESS THIS:
'^M{0-3)$'
(A|B) - MATCHES A OR B BUT NOT BOTH
you should never “chain” the search() and groups() methods in production code. If the search() method returns no matches, it returns None, not a regular expression match object. Calling None.groups() raises a perfectly obvious exception: None doesn’t have a groups() method. (Of course, it’s slightly less obvious when you get this exception from deep within your code. Yes, I speak from experience here.)
55. REGEX USE CASE :
<html lang="en" dir="ltr" class="client-nojs">
<head>
<meta charset="UTF-8" />
<title>Guido van Rossum - Wikipedia, the free encyclopedia</title>
<script>document.documentElement.className = document.documentElement.className.replace( /(^|\s)client-nojs(\s|$)/, "$1client-js$2" );</script>
SAY YOU WISH TO GET ALL THE TAGS ELEMENTS.
<.*> - * means 1 or more. * is greedy by default. SO, it will start at the first < and gobble as much as possible - here,the entire thing before matching the last >
to make it non-greedy ; that is gobble as little as possible :
<.*?> - this gets us the tags
also valid regex : <.+?> - "+" matches 0 or more characets, but ? forces it to gobble as little as possible.
The square brackets mean “match exactly one of these characters.”
>>> re.sub('[abc]', 'o', 'caps') ④
'oops'
re.sub replaces all of the matches, not just the first one. So this regular expression turns caps into oops, because both the c and the a get turned into o.
>>> re.sub('([^aeiou])y$', r'\1ies', 'vacancy') ② - here, `cy` matches. So, when replacing : replace group 1 by itself. IE [^aeiou] by itself. and `y` by ies.
'vacancies'
56. HOW TO OPEN FILES
with open('plural4-rules.txt', encoding='utf-8') as pattern_file: ②
for line in pattern_file: ③
print line
############################
EXPERIMENTATION
import re
def plural(noun):
if re.search('[sxz]$', noun): ①
return re.sub('$', 'es', noun) ②
elif re.search('[^aeioudgkprt]h$', noun):
return re.sub('$', 'es', noun)
elif re.search('[^aeiou]y$', noun):
return re.sub('y$', 'ies', noun)
else:
return noun + 's'
ANOTHER WAY :
import re
def match_sxz(noun):
return re.search('[sxz]$', noun)
def apply_sxz(noun):
return re.sub('$', 'es', noun)
def match_h(noun):
return re.search('[^aeioudgkprt]h$', noun)
def apply_h(noun):
return re.sub('$', 'es', noun)
def match_y(noun): ①
return re.search('[^aeiou]y$', noun)